18 August 2026Practice
Lessons ledger: the reviewer we never checked
What is this ledger?
A running record of things that went wrong in our own work, what each one cost us, and the rule we now follow because of it. Every entry keeps the same shape: what happened, what it cost, what rule came out of it, and where that rule falls short.
Two limits govern every entry, this one and all the ones to come. First, an incident appears here only where the loss was ours alone. Nothing is drawn from a client’s system or a client’s work, and nothing is anonymised to sneak past that line; where the strongest example belongs to a client, we use the next strongest that does not, and this page is exactly that. Second, one incident is one incident. A ledger entry is evidence that something can happen, never proof of how often it does.
What happened?
For a week in July 2026, one of our review agents claimed an independence that nobody had verified.
Some context, briefly. Our changes are reviewed before they merge, and the author is never the reviewer. Part of that reviewing is done by an isolated agent: it is handed the change and the project’s rules, and it is supposed to be sealed, with no network, no live systems, and no credential for the code it is judging, so that its verdict cannot be influenced by the thing under review.
The agent’s own instructions described that seal as a verified fact rather than a design goal. Nobody had checked, for a week, across every review the agent ran.
The first time we made the agent check, instead of telling it what to believe, the check took under a minute. Its opening move was to list the tools and access it actually held, and in that list was a live, working credential for the very repository it was reviewing, carrying the right to merge code into it. Across the next seven runs, the same opening check found four more real gaps in the seal. Not one of them was found by reasoning about the configuration. Every one was found by the agent measuring itself.
The counts in this entry, the runs and the gaps alike, come from our own working record of that week, and that record is not public, so you are taking them on our word rather than being able to check them.
It was not a one-off flaw in one agent, either. In the same stretch of work, a documented option for restricting an agent’s tools was passed on two consecutive runs and did nothing at all. The only reason anyone noticed is that the agent was now reporting what it actually held.
What did it cost?
A week of reviews whose central claim we could not stand behind. The verdicts themselves may all have been sound, and that is exactly the problem: we had no way to tell. A review is worth something because the reviewer is separate from the work, and for that week the separation was an assertion, not a fact. The honest description of every one of those reviews is “probably fine”, which is not what an independent review is for.
It also cost the work of closing what the new check went on to find, five real gaps in all. And it cost a comfortable belief: that a restriction we had written down was a restriction that existed.
What rule came out of it?
A constrained thing proves its constraint, as its first action, from direct evidence, on every run. Our sealed reviewers now begin by listing what they can actually see and touch, and each report opens with that list. Where the agent cannot establish something, it says unestablished, never absent. An agent reporting a clean seal it did not check is worse than no agent at all, because the report reads as assurance.
And the sibling rule, from the option that did nothing: a restriction is unverified until something observes it taking effect. A setting you applied is a hope. A setting whose effect you measured is a control.
Where does the rule fall short?
It is not free, and it can decay into ritual. The opening check runs on every review, and almost every time it finds nothing, which is exactly the condition under which people stop reading the result. A check strict enough to fail every run gets waved through within a couple of runs and then catches nothing, so ours has to distinguish a breach that poisons the work, which stops it, from one that can be declared and worked around. Drawing that line is judgement, and we can get it wrong in both directions.
The rule also covers only what an agent can observe about itself. A gap the agent has no way to see stays open, and no self-report will ever mention it. This check moved our confidence from “configured” to “measured”, not to “certain”, and those are different words on purpose.
What would change this?
A breach this check misses. If a seal fails in a way the opening self-check cannot see, and we find that out some other way, this entry gets a successor and the rule gets a revision. The ledger is written to be corrected in the open rather than quietly.
The other thing that would change it is better enforcement. Today the seal is held by discipline and verified by measurement, because the platforms these agents run on do not always honour a strict allow-list of capabilities. The day the boundary can be built rather than checked, the rule above becomes a backstop instead of the control, and we would say so here.