The codebase is a big ball of mud
Aspects: effective action (meta-systematic) · epistemology (relating formal and informal) · contingencies (messes to manage)
Signals you are in this situation
Section titled “Signals you are in this situation”- Nobody can say what a change will affect without trying it.
- Simple fixes take days and break unrelated areas.
- Tests pass and production still surprises.
- There is a rewrite plan, and there has been one for years.
- The knowledge of how it really works lives in two people.
The situation
Section titled “The situation”Will Larson describes one of the smartest engineers he ever worked with, who could not finish projects. Every task exposed a related problem, then a surrounding one, until the original change had become insurmountable. Years later, trying to extend large existing systems himself, Larson recognised the experience. The systems appeared to have properties, such as all writes going through one log, but each property had a few caveats: a performance shortcut here, an exception two products made to hit a launch date there. Persistent violations void an abstraction, and once enough craters have accumulated no properties remain, so any abstract model of the system’s behaviour is wrong in ways you cannot predict.
Brian Foote and Joseph Yoder gave this kind of system its name in 1997 and called it the de facto standard architecture: haphazardly structured, sprawling, held together with expedient repairs, with information shared promiscuously among distant parts. Their point was not to condemn it. Big balls of mud endure because they work, and because the forces that produce them are real: deadlines, changing requirements, code that was meant to be thrown away and was not, and the sound judgment that a quick entry into the market can be worth more than an elaborate architecture built on unproven guesses.
The meta-rational perspective
Section titled “The meta-rational perspective”A rational approach to modifying code starts from a model of the design and reasons forward: if the system has these properties, then this change will have these effects. That approach requires the properties to hold. In a big ball of mud they do not, and Larson’s observation is that building a more nuanced model does not help. You refine it into rich sophistication and it is still wrong. Nelson Elhage draws the same conclusion from the infrastructure side: such systems work, insofar as they work, through specific pairwise interactions between distant parts, not through any comprehensible core logic, so reading the source cannot answer the questions you need answered.
What works instead is a switch from abstract to empirical reasoning. Replace assertions about properties with observed behaviour. Ask the running system questions through logs, metrics, traces, and targeted instrumentation. Document what it does rather than what it was intended to do. Mental models are still useful for forming hypotheses, but every hypothesis has to be checked against the territory before it is relied on. This is the epistemological row of Chapman’s table in practice: the formal description and the informal reality have drifted apart, and the work is to relate them again rather than to trust either alone.
The second meta-rational move is to stop treating the mud as a single problem. Kevin Simler’s image is of a codebase as an organism that decays whenever anything around it changes, and stays healthy only through dozens of small interventions. Matt Belcher lists the signs that those interventions have stopped: changes that break unrelated areas, feature delivery that slows, a domain model that no longer fits the business, and, tellingly, large refactorings, which are what you get when small ones were skipped for too long. Foote and Yoder are blunt about the alternative. A system a business depends on cannot be taken down for an overhaul, and after a big-bang change it is impossible to tell which of many modifications caused the new failure. The way back is the way Larson describes: pick one property to reassert, observe the behaviour around it, run a pilot, migrate into it, and repeat. Each step makes the next one easier. Geoffrey Litt’s reading of Stewart Brand adds why this is not a defeat: buildings that survive are the ones whose layers can change at different rates, and code that is left rough enough to be altered by its inhabitants may be more adaptable than code that was designed to be admired.
Failure path
Section titled “Failure path”The developer reasons from what the system was designed to do, does not verify that the design still holds, and makes the locally easy change rather than the design change the request actually needs. The proper fix is deferred to a rewrite that never arrives. The existing tests pass, because they encode the same stale model. Then a distant part of the system, which depended on the old behaviour through one of the caveats nobody knew about, breaks, and the patch that fixes it adds one more exception to the pile.
Corrected path
Section titled “Corrected path”The request still lands in the same mud. The difference is that the first step is observation rather than modelling, the vocabulary the mud has blurred is rebuilt before anything is designed, and the change is scoped to a single property that can be made true at one seam and fenced with tests before the code moves. The loop from step 6 goes back to observation, not to the model, because a surprise in a property-less system is information about the territory. The loop from step 7 is the migration Larson describes: the next property, a little easier than the last.
Skills for this situation
Section titled “Skills for this situation”| Skill | When to invoke it here | What it changes | Example |
|---|---|---|---|
research |
Before touching code, when the question is what the system actually does. | Sends a background agent to read the real sources, including logs, metrics, and the code paths that run in production, and to write down observed behaviour with citations rather than intended behaviour. | /research what actually happens when an order is cancelled after fulfilment starts; cite log lines and code paths |
diagnosing-bugs |
When an observation contradicts the model, even if nothing is formally broken. | Its loop of reproduce, hypothesise, instrument, and isolate is the empirical method Larson and Elhage recommend, applied to a property rather than a defect. | /diagnose why some writes bypass the event log; treat it as a bug in the invariant |
domain-modeling |
Once observation shows one term covering several behaviours. | Rebuilds the glossary the mud lost. Names the distinctions the code has been blurring, records them in CONTEXT.md, and captures the property you intend to reassert as an ADR. |
/domain-modeling: the codebase uses "order" for three different lifecycle objects; separate them |
codebase-design and improve-codebase-architecture |
When choosing which property to reassert first and where. | Gives the vocabulary of deep modules and seams, and a scan that ranks deepening opportunities so the first migration is the one with the most leverage, not the one nearest to hand. | /improve-codebase-architecture |
tdd |
Before the change, at the chosen seam. | Characterisation tests that pin observed behaviour, so the migration can proceed in small steps with each step rejected immediately if it changes anything it should not. | /tdd characterisation tests for the cancellation path as it behaves today |
implement and code-review |
For the migration itself. | Small steps, typechecks and tests on every one, and a review against the ADR that states the property being reasserted rather than against the original request. | /code-review against main: does this hold the invariant recorded in ADR-0004? |
to-tickets |
At the end of each migration. | Turns the exceptions you found but did not fix into tickets with blocking edges, so the next property is queued rather than forgotten. | /to-tickets from the remaining caveats found while reasserting the event-log invariant |
Sources
Section titled “Sources”- Brian Foote and Joseph Yoder, Big Ball of Mud, 1997.
- Will Larson, You can’t reason about big balls of mud.
- Nelson Elhage, Systems that defy detailed understanding.
- Kevin Simler, A Codebase is an Organism.
- Matt Belcher, Signs Your Software is Rotting.
- Geoffrey Litt, thread on How Buildings Learn, after Stewart Brand.
- David Chapman, Meta-rational software development: Readings, section on the nebulosity of software itself.