← Writing

Docs are a wish. Review does not scale. So the rules run.

8 min readArchitectureTestingEngineering Practice
Contents· The two ways a decision dies

A decision gets made in a pull request. Six months later the person who made it has moved on, the reasoning is three hundred comments deep, and someone deletes it in a tidy-up. I stopped trying to solve that with documentation. Each decision is a test file now, one decision per file, and it fails the build if somebody undoes it. There are 855 of them.

The two ways a decision dies

Every architectural decision you make has to survive two things: people leaving, and people tidying up. The second one is worse, because it comes with good intentions.

The usual answers both fail. Documentation is a wish. It is correct on the day it is written and drifts from then on, and nothing tells you when it has stopped being true. Code review works exactly as long as the reviewer remembers. It does not scale past the few decisions any one person holds in their head, and it does not survive that person at all.

Both share a flaw. They describe the rule somewhere the rule is not enforced. The gap between the two is where the decision quietly dies.

What a rule looks like when it runs

Here is a real one, from Poplitu, the forms product I am building. Several parts of it send email, and metering was added so customers could be billed for what they use. The decision underneath it was that metering must never become a limit on transactional mail, because a free tier of zero would stop a form replying to the person who just filled it in.

That decision is not a comment. It is a file, and the file explains itself:

METERING IS AN ACCOUNT, AND IT MUST NOT BECOME A CAP.

This builds the number a reader would reach for, and puts it
within easy reach of everything that sends. Nothing in the code
says it may not be read as a limit, so this file does.
The opening of that rule’s docblock, abridged. The file sits beside the code it guards.

Read the last sentence again, because it is the whole idea. Nothing in the code says it may not be consulted, so this file does. The test does not check that the code works. It checks that a specific mistake has not been made. The mistake is plausible, it looks like a fix, and the person who makes it will be trying to help.

The one that caught false product copy

This is the example that changed how I think about what a test is for.

The product connects to a long list of third-party services. Four separate places in the interface explained to the user how connecting works, and all four described an OAuth flow, the pop-up, the permissions screen, the redirect back.

A good share of those providers do not use OAuth at all. For those, the explanation was not vague or generic. It was false, in four places, and no test could fail because nothing about it was broken. The code did exactly what it was written to do.

Tests check that code works.

They can also check that a claim you make to a user is still true

The rule that exists now says no surface may explain connecting on its own account. It walks the interface, finds any copy describing a connection flow, and fails if it is not derived from what the provider actually does. Product copy, held to the same standard as a type signature.

Four kinds of rule

They are not all the same weight, and pretending they are would be the easiest way to oversell this. Sorted honestly, they fall into four groups.

KindWhat it holdsExampleHonest weight
HygieneConventions a linter cannot expresstests are typechecked, file length capsLow. Useful, not clever.
ContractTwo sides of a boundary agreeingpublished types match the generated clientMedium. Catches real breakage.
BoundaryA security or tenancy lineevery table holding customer data is covered by a policyHigh. Cheapest place to catch these.
DecisionA choice someone will try to undometering must not become a capHighest. Nothing else holds these.
Counts are from one command, shown at the end. The fourth group is the one people mean when they say “architecture test”. The first is closer to a lint rule and I would not defend it as more than that.

The last group is the reason to do any of this. The other three are things you could get from better tooling. A decision rule is different: it encodes a judgment, including the reasoning, in the one place a future engineer cannot avoid reading, the failure output of a build they need to pass.

The name is the claim

Early ones were named after the thing they inspected. body-validation, and a dozen like it. Fine, and forgettable.

The newer ones are sentences. a-dropped-table-was-retired-first. a-crm-write-goes-through-one-door. There are 146 named that way now; the other 707 still carry the name of the thing they inspect.

This sounds cosmetic and is not. When a test named body-validation fails, you learn a test failed. When a-dropped-table-was-retired-first fails, you learn that a migration is about to drop a table something is still reading from, which is a sentence about your product that somebody needs to read. The failure output is the documentation, and it cannot drift, because it is produced by the thing it describes.

Rules that argue with each other

The ones I am most pleased with cite each other. One gate holds every file in a directory to a rule. A second gate covers a neighbouring directory, and its docblock argues why it is a second gate rather than the first one widened, quoting the first gate’s own reasoning about why it sweeps a directory instead of naming three files.

That is an architecture conversation, conducted between two files, that runs on every commit. No wiki page does that.

What 855 tests actually cost

Three real costs, and I would not trust this piece if it skipped them.

They run with everything else

These are ordinary test files. They are not a separate gate, they have no special runner, and they add to the time every commit takes. That is the price, and it is the right price , a rule that runs somewhere optional is a rule that gets skipped.

A vacuous rule is worse than no rule

A test that asserts nothing still passes, and now you believe something is protected when it is not. This is the real failure mode, and it is easy to hit, a selector that matches nothing, a scan of a directory that has since moved.

The discipline is dull and non-negotiable: after writing the rule, break the thing on purpose and watch it fail. If reverting the fix does not turn the test red, the test is decoration. Several of mine were, until they were checked.

Not all 855 are worth 855

Two thirds of them live in one service. Some are close to lint rules wearing a test costume. I would defend the decision and boundary groups to anyone; the hygiene group is convenience, and calling it architecture would be flattering myself.

So if you want the figure the argument actually rests on, it is not 855. It is 146, the ones named as a claim rather than after the thing they inspect. 855 is a count of files. 146 is a count of claims somebody wrote down and now has to keep true, and that is the number I would want to be judged on.

When this is the wrong idea

Do not do this on a product still deciding what it is. An invariant freezes a decision, and early on you want decisions cheap to reverse. Freezing them is precisely wrong.

It starts paying when one of these is true. The codebase outgrew the number of decisions one person can hold. More than one person changes it. Or the cost of a specific mistake is high enough that catching it late is unacceptable, tenancy leaks, silent data loss, a compliance claim made to a customer.

Until then, a comment is fine. The rule is for when a comment has stopped being enough.

Start with one

Not 855. One. Pick the decision on your current project that would do the most damage if it were quietly undone, the one you would explain in a code review if you happened to be looking.

Then write it down as a test that fails, in four parts:

  1. Name the file as the claim, so the failure is a sentence.
  2. Write the reasoning in the docblock, including what the tempting wrong fix looks like.
  3. Assert the structure, not the behaviour.
  4. Break it on purpose. Watch it go red. Put it back.

You will know within a month whether it earned its place, because either it caught something or it sat there. Mine caught a screen describing a connection flow that a good share of the providers do not use, which no amount of code review had noticed, because nothing was broken.

Counting these yourself

You cannot check my count, and you should not take it on faith. Run it against your own repository instead. One rule is one file, so counting them is counting files, and the useful number is not mine, it is how few come back:

find . -name "*.invariant.spec.ts" -o -name "*.invariant.test.ts" \
  | grep -v node_modules | wc -l
855 files when this was written, 146 of them named as a claim. The first was created on 14 June 2026; the most recent while this piece was being drafted.