Stop the Line
Repair known defects within scope; stop at missing authority, unresolved intent or spent limits.
Caller-owned acceptance defines the stopping boundaries. Factor XII includes failures, stopped runs and repair costs when deciding what the next cycle should change.
The rule
Stop when the agent lacks authority, intent is unresolved or real limits are exhausted.
Repair known defects inside accepted scope while meaningful progress remains possible. A genuine stall calls for diagnosing the assumption that failed, rather than restarting the same attempt.
An honest incomplete outcome preserves what was checked and what remains unknown. It does not count as completed capability. A new context, helper or retry cannot renew spent limits or expand authority.
Agents often miss that a request is incomplete, so set the stop conditions before the run and tell the agent to ask. When a session has collected failed attempts, carry what was learned into a new brief and start clean. The spent limits still count.
The evidencemeasured
Experiments test the two moves in this rule: asking when intent is unclear, and stopping to escalate when a task cannot be finished honestly. The rule on spent limits is operating practice.
- Ambig-SWE, 2025 removed details from software tasks. Models struggled to tell a well-specified request from an underspecified one. When they did interact, performance improved by up to 74% over the non-interactive setting.
- ImpossibleBench, 2025 gave agents a way to flag an impossible task for a person and stop. GPT-5's rate of passing by breaking the rules fell from 54% to 9%, and o3's from 49% to 12%. The effect was much smaller for Claude Opus 4.1.
- Gomez, 2026 tested an escalation tool with a stated policy on 8 frontier models. Reward hacking fell from 23.6% to 5.3%, with no detectable cost or performance overhead.
- In the MAST taxonomy of multi-agent failures, step repetition was the most common mode at 15.7%, and being unaware of termination conditions accounted for 12.4%.
Starting clean after repeated failed attempts is vendor guidance, from Anthropic's Claude Code documentation, and fits the context studies cited in Factor I.
Put it to work
- Set stop conditions before the run and tell the agent to ask when intent is unclear.
- Repair a known check failure directly within the accepted scope.
- Stop before publication or execution that needs missing approval.
- Diagnose recurring failures before attempting another approach.
- Leave a compact handoff with current output, evidence, gaps and the needed decision.
Example · the monthly operations report
The report's ledger totals disagree with the source owner. Investigate a known parsing defect within scope. If source access is missing or the discrepancy cannot be resolved within the limit, return the draft and failed reconciliation instead of publishing an invented total.
Failure signal
- Retries repeat the same failure without new evidence.
- The agent silently changes intent or permissions to finish.
- A blocked result is presented as an accepted report.
Done looks like
Work either meets its required acceptance or stops with a truthful, recoverable handoff. The caller can see what decision or evidence is needed next.