Factor VII of XIIGate the Work

No Verdict, Not Done

Check the exact output against intent; use fresh independent judgment where risk calls for it.

In the loop

Execution produces an exact candidate output. Checks and any required independent judgment establish whether it meets intent, before Factor VIII preserves accepted behavior.

The rule

Do not equate the agent's “done” with an accepted result.

Run the checks that establish the requested behavior. Ordinary work can finish on those checks under the caller's policy; software work may also require CI.

Use fresh, author-distinct judgment when requested, when a mistake is hard to undo or when deterministic checks do not cover the result. That reviewer examines the exact output against unchanged acceptance criteria. PASS needs evidence for every criterion; FAIL names a demonstrated failure; NOT_PROVEN names missing evidence. A verdict does not grant permission to publish.

A fresh context removes the author's reasoning from the review. It leaves one bias in place: models tend to favor text they recognize as their own. Where a mistake costs the most, the caller can add a reviewer from a different model family.

The evidencemeasured

Studies compare what agents report, what tests accept and what people accept.

  • METR, 2026 had four maintainers review 296 agent-written pull requests for SWE-bench Verified issues. By their judgment, roughly half of the pull requests that passed the tests would not be merged.
  • Wang et al., 2025 retested patches from the same benchmark. 7.8% counted as correct while failing the full developer-written test suite, and 29.6% of passing patches behaved differently from the human fix.
  • Reddy et al., 2026 ran 1,980 code-modernization calls on 11 production models, then asked each model whether its own output preserved behavior. The authoring model approved 31.7% of the outputs that had changed behavior. Miss rates ran from 0% on five models to 100% on one. A 2024 survey of earlier work found self-correction succeeded when reliable external feedback was available.
  • Panickssery et al., 2024 found LLM evaluators can tell their own output from other text, and the better a model recognizes its own text, the more it favors it.

METR notes its agents had no chance to respond to review as a human author would. Repairing findings and rerunning the affected checks is that step.

Independent review has limits too. In ImpossibleBench, 2025, LLM monitors caught 86 to 89% of rule-breaking on simpler tasks and 42 to 65% on complex ones. None of these studies measures how much a fresh reviewer adds on real work.

Put it to work

  • Bind checks and judgment to the version actually being accepted.
  • Separate check facts from conclusions about meaning.
  • Use one fresh review where risk calls for it.
  • Repair findings within scope and rerun affected checks; further review follows caller policy.

Example · the monthly operations report

Reconciliation and citation checks cover the report's totals and references. A fresh reviewer checks a consequential recommendation against the cited evidence. Missing support remains NOT_PROVEN even when the arithmetic passes.

Failure signal

  • “Done” relies only on the author's confidence.
  • Review covers an earlier version or changed criteria.
  • A missing criterion disappears from the completion claim.

Done looks like

The completion claim names the exact output, checked criteria and any gaps. Required judgment is independent; the caller retains acceptance and publication authority.