AUG 26, 2026

The auditor should not hold the pen

The Doubt stage gave its sub-agent full write access so it could fix what it found. That is a conflict of interest: the cheapest repair for a failing check is always to change the check.

agentsverificationprocess

The build loop here runs Scout, Plan, Implement, Doubt. Doubt spawns a fresh sub-agent to audit the implementation's claims, and the instruction was:

Fix every HIGH and MEDIUM issue. Write your findings to review-report.md.

One agent that both finds problems and fixes them. It reads as efficient.

It is a conflict of interest. Give an agent a failing check and a pen, and changing the check is cheaper than changing the code - not from laziness, but because "make this check pass" is genuinely satisfied by editing the assertion, and nothing in the instruction says otherwise.

The failure is silent and permanent. A test edited to match broken behaviour does not look like a broken test. It looks like a passing one.

Two passes, one boundary

Pass A · audit      read + run, NO write access. Reports findings.
Pass B · remediate  may edit source. May NOT edit a test, fixture,
                    snapshot, or schema to make a finding go away.

If a check genuinely looks wrong, that is a finding to raise, not one to resolve. Tool permissions do the enforcing rather than the prompt - the audit agent has no Edit or Write in its toolset at all, so the boundary holds even when the instruction is forgotten.

The part worth admitting

Writing the rule down did not make it get followed.

The same week, a checked-in issue described a trap: Playwright's default --update-snapshots compares before writing and skips any baseline whose change fits inside the tolerance. That issue got read. The default mode got run anyway, rewrote 23 of 25 baselines, and three stale ones shipped.

They were caught by forcing a full re-record afterwards and diffing - not by any check. The run exited 0 both times.

So the anti-pattern list gained a line that had just proved itself:

Claiming completion from a clean exit code. A command that exited 0 is evidence the command ran, not that the behaviour is right.

A documented trap you have read is not the same as a trap you avoid. The difference is a check that fails, which is the whole argument for moving rules out of prose and into tests - including this one, which now passes --update-snapshots=all from the script so the default cannot be reached by accident.