Factor XII of XIIGovern the Loop

Price the Proof

Measure cost per accepted output, include failures and repairs, and feed findings into the next brief.

In the loop

Count the cost of accepted outputs, repairs and stopped runs. Feed useful findings into the next brief in Factor I, closing the loop.

The rule

Measure the cost of work that meets acceptance, including checks, review and rework.

Compare similar outcomes under the same criteria. Fast drafts are useful, but draft speed alone does not tell you what reliable delivery costs.

Include failure and recurrence alongside time or spend. An empty green report, weaker test or omitted criterion cannot count as an efficiency gain. Improve the method only where evidence identifies avoidable work; keep protection proportionate to risk.

Record the numbers instead of trusting how fast the work felt. Agents vary from run to run, so judge a method over several runs before trusting its pass rate or its cost.

The evidencemeasured

A randomized trial and cost-controlled benchmarks test this rule directly.

  • METR, 2025 randomly allowed or disallowed AI tools on 246 tasks done by 16 experienced open-source developers. With AI allowed, tasks took 19% longer. Afterward the developers estimated AI had made them 20% faster.
  • METR's 2026 follow-up could not produce a clean number, because many developers would not submit tasks they might have to do without AI. METR believes developers are faster with AI now than in early 2025 and is redesigning the study.
  • Kapoor et al., 2024 showed that judging agents on accuracy alone produced needlessly costly systems, and that cost could be cut sharply at the same accuracy.
  • τ-bench, 2024 ran each task repeatedly. The best agents succeeded on fewer than half the tasks in a single try, and on all eight tries for fewer than a quarter of retail tasks.
  • DORA, 2024 found AI adoption raised individual productivity while lowering software delivery stability and throughput.

Self-estimates miss and even a careful trial can break, so keep your own numbers on your own work.

Put it to work

  • State the outcome and denominator behind each metric.
  • Include failed attempts, review, repairs and operator time.
  • Pair speed or cost with acceptance and recurrence measures.
  • Run a method more than once before trusting its pass rate or cost.
  • Use observed friction to improve the next brief, source access or paved path.

Example · the monthly operations report

Compare monthly reports by cost per accepted report, including reconciliation and review. Track unresolved discrepancies and repeated omissions too. If locating the approved ledger caused most of the delay, put its current source link in the next brief; do not remove reconciliation to make the run faster.

Failure signal

  • Cost excludes failed runs or human repair time.
  • Throughput rises while omissions return.
  • A benchmark changes acceptance between comparisons.

Done looks like

The caller can see what reliable outputs cost and which supported change should improve the next cycle. That feedback becomes the next brief.