DevOps discipline for all agentic work.
Twelve factors for agent output you can check: clear intent, bounded execution, checked results and memory for the next task.
- Totals reconcile with the approved ledgercheck passed
- Every claim cites a sourcecheck passed
- Personal details excludedcheck passed
- Recommendation supported by cited evidenceno evidence
The arithmetic passes. One criterion has no evidence, so the report is not done.
“Done” fails in three ways.
Judgment
The output passes narrow checks but an unsupported conclusion still misses the accepted request.
The fix. Check the output against accepted intent; use fresh judgment where risk calls for it.
Durable Context
The next task repeats an earlier mistake because its decision and supporting evidence cannot be retrieved.
The fix. Linked work history and memory keep evidence available for later tasks.
Loop Closure
The output is checked, but findings never reach the next brief or prevent a repeated failure.
The fix. Confirm repairs and carry supported feedback into the next brief.
What the studies measured.
Each factor essay cites the research behind its rule and says where a rule is practice that has not been tested.
of test-passing agent pull requests would not be merged, four maintainers judged.
is what developers believed AI made them in a 2025 trial. Measured, they were 19% slower.
is the range for multi-agent systems against one agent, depending on the shape of the task.
is how far GPT-5’s rate of passing impossible tasks by breaking the rules fell once it could flag the task and stop.
attack success against most of 12 published defenses for jailbreaks and prompt injection.
Four tiers. One loop.
Brief, rail, gate, govern, then the next brief. Operating principles applied in proportion to the task, not twelve mandatory steps.
The context that wrote a change can’t issue its own PASS.
- 01
Accepted intent
The caller’s outcome, examples, scope and authority.
- 02
Implementation and checks
Native work, with the checks that establish the requested behavior.
- 03
One fresh read
Where a mistake is costly, an author-distinct reviewer judges the exact result.
- 04
Finish
The caller’s policy controls acceptance and delivery.
In AgentOps v3.9.0: Accepted intent → native implementation and checks → one fresh read where a mistake is costly → finish
Pick the row that matches your task.
Each entry point is an independent choice. A clear task can go straight to your agent.
Not sure where to start? Try the read-only Research task →
AgentOps puts the factors on paved paths.
Its optional skills and the ao CLI work in the coding agent you already use. Start with a read-only Research task in a repository you know. This site tracks AgentOps v3.9.0.
claude plugin marketplace add boshu2/agentops
claude plugin install agentops@agentops-marketplaceCodex and Cursor commands are on the install page.
Your agent said done. Something else proved it. The proof is yours.