Skip to content

The obvious data model breaks on reality

Aspects: ontology · categories (reflection on boundaries) · contingencies (anomalous)

  • Validation keeps rejecting real records: names, addresses, dates, quantities.
  • Every sprint adds another special case to the same table or type.
  • A field is nullable for now and nobody can say what null means.
  • Support tickets describe the same edge case in different words.
  • A column called notes or extra is doing structured work.

Hillel Wayne opens a confectionery textbook at random and sketches a data model for its recipes: a recipe is a list of ingredients, and an ingredient is a food with a mass. The next page calls for twenty-five lemon skins, so a measure has to be a mass or a count. The page after that lists foil cups, which are not food. Then an ingredient turns out to be optional, then one is a choice between two foods, then a recipe contains a sub-recipe with its own list. Each page that seemed to complete the model breaks it, and the model that survives is more complex, harder to understand, and more fragile than the one he started with.

The same thing happens with people. John Graham-Cumming’s surname contains a hyphen, and a web form rejected it with a message saying his name contained invalid characters; every airline he has flown with has recorded him as GRAHAMCUMMING. The lists collected under the heading of falsehoods programmers believe do this systematically for names, addresses, email, time, and much else: each entry is an obvious assumption that reality violates.

William Kent’s book on data modelling starts from the observation that data structures are maps, and that no map is the territory. An information system is a model of a small part of the world, built on the expectation that one construct inside it corresponds to one thing outside it, and even that expectation fails at once. Is a part one kind of thing, as in an inventory file, or one physical object, as in a quality-control file? Each application resolves the ambiguity silently, from context, and the resolution is lost the moment the two files are integrated. Kent’s conjecture is that information is too amorphous and too subjective to be pinned down completely by any formal model, and that models are nevertheless indispensable.

Chapman’s account of ontological remodelling gives the constructive side. Formal rationality works inside a fixed ontology, usually implicit, and that is fine while the ontology is good enough for the job. When it is not, rationality has no way to repair the breach, because it cannot see its own categories. Meta-rationality treats the categories as malleable: a category can be dropped, demoted to an informal heuristic, or given a new formal meaning, and the choice is made by reflecting on the boundaries. Wayne’s practical proposal is the wastebasket: a deliberately unstructured place for the cases that will not fit, such as a free-text field, with the structured model reserved for the cases that matter. Which cases those are cannot be settled by looking at the data. It depends on what the software is for and who will pay for each mismatch.

The failure is not that the schema was wrong. Every schema is wrong somewhere. The failure is a process that treats each anomaly as a bug to be patched in place, so that the model accretes special cases without anyone ever deciding where its boundary should be.

1 Look at typical cases the first three recipes 2 Fix the schema ingredient = food + mass 3 Validate strictly reject what does not fit 4 Build on the schema forms, storage, reports 5 Ship tests pass on the typical cases 6 Real data arrives 25 skins, foil cups, a hyphenated name 7 Patch per case special case #14 and counting each exception becomes a new column and a new bug
How the situation goes wrong when handled purely rationallyHow to read

The schema is drawn from the typical cases, fixed, and enforced with strict validation. Everything downstream is built on it. When real data arrives with foil cups, counts instead of masses, and hyphenated names, the validation rejects it, and each rejection is handled as a special case. The special cases accumulate into exactly the complex, fragile model Wayne warned about, with the additional cost that the users whose data was rejected have been told that their names are invalid.

1 Sample widely real data plus the falsehood lists research: Sends a background agent to primary sources and writes cited findings to a file.research 2 Model the edges invent the odd cases on purpose domain-modeling: Challenges overloaded terms and writes the glossary and decision records.domain-modeling 3 Draw the boundary what gets structure, what gets a wastebasket grill-with-docs: Grilling that also writes glossary entries and ADRs as decisions settle.grill-with-docs ADRs: Architecture decision records: dated, reasoned decisions kept in docs/adr/.ADRs 4 Test the boundary one test per falsehood tdd: The red-green-refactor loop with rules for tests worth keeping.tdd 5 Reuse what exists a proven library for names, addresses, time research: Sends a background agent to primary sources and writes cited findings to a file.research 6 Ship with a way out free text where the model gives up implement: Implements a spec or tickets test-first at agreed seams.implement 7 Review anomalies remodel, do not patch domain-modeling: Challenges overloaded terms and writes the glossary and decision records.domain-modeling a real anomaly: revise the categories, not the special cases a test would not pass: move the boundary
The same path with skills inserted at the break pointsHow to read

The shape is the same: a model is still drawn, built on, and shipped. The difference is that edge cases are hunted before the schema is fixed, the boundary between structured and unstructured is a recorded decision rather than an accident, tests encode the falsehoods that were considered, and established libraries are preferred for the domains where the falsehood lists are longest. The loops route anomalies back to the categories rather than into the patch queue.

Skill When to invoke it here What it changes Example
research Before drawing the schema, and again before writing a parser. A background agent reads the falsehood lists and the primary specifications for the domain, and records which established library already handles it. Names, addresses, email, and time are all better bought than built. /research which falsehoods about postal addresses apply to us, and which library handles them
domain-modeling While the model is still a sketch. Invents edge-case scenarios on purpose, challenges each category, and writes down what an ingredient, a measure, or a name is for this system. Records the result in CONTEXT.md. /domain-modeling what counts as an ingredient here: foil cups, optional items, either-or choices, sub-recipes?
grill-with-docs When deciding what the model will refuse to represent. Interrogates the boundary until it is a decision: which cases get structure, which go to a wastebasket field, and who bears the cost of each. Writes the decision as an ADR. /grill-with-docs where the recipe model stops and free text begins
tdd Once the boundary is decided. Turns each considered falsehood into a test, so the boundary is enforced by tests rather than by validation rules that reject people. A test that cannot pass sends the boundary back for revision. /tdd a name field that accepts hyphens, apostrophes, single names, and four-part names
implement For the build. Builds the model with its escape hatch in place, so an unforeseen case degrades to unstructured data rather than to a rejection. /implement the recipe model from the ADR, with the free-text fallback
domain-modeling, again Whenever a new anomaly arrives in production. Treats the anomaly as evidence about the categories. Sometimes a category should be split, sometimes demoted to informal status, sometimes the wastebasket is the right answer. What it should not become is special case fourteen. /domain-modeling a recipe arrived that lists an ingredient by volume; revisit the measure category