DevOps discipline for all agentic work.

Twelve factors for agent output you can check: clear intent, bounded execution, checked results and memory for the next task.

validatemonthly-ops-report.md
  • Totals reconcile with the approved ledgercheck passed
  • Every claim cites a sourcecheck passed
  • Personal details excludedcheck passed
  • Recommendation supported by cited evidenceno evidence
NOT_PROVEN

The arithmetic passes. One criterion has no evidence, so the report is not done.

An illustration of the example every factor essay follows.
The problem

“Done” fails in three ways.

Judgment

The output passes narrow checks but an unsupported conclusion still misses the accepted request.

The fix. Check the output against accepted intent; use fresh judgment where risk calls for it.

Durable Context

The next task repeats an earlier mistake because its decision and supporting evidence cannot be retrieved.

The fix. Linked work history and memory keep evidence available for later tasks.

Loop Closure

The output is checked, but findings never reach the next brief or prevent a repeated failure.

The fix. Confirm repairs and carry supported feedback into the next brief.

The route to done

The context that wrote a change can’t issue its own PASS.

  1. 01

    Accepted intent

    The caller’s outcome, examples, scope and authority.

  2. 02

    Implementation and checks

    Native work, with the checks that establish the requested behavior.

  3. 03

    One fresh read

    Where a mistake is costly, an author-distinct reviewer judges the exact result.

  4. 04

    Finish

    The caller’s policy controls acceptance and delivery.

A fresh reviewer returns one of three verdicts:PASS evidence for every criterionFAIL a demonstrated failureNOT_PROVEN the missing evidence, named

In AgentOps v3.9.0: Accepted intent → native implementation and checks → one fresh read where a mistake is costly → finish

Start where the work is

Pick the row that matches your task.

Each entry point is an independent choice. A clear task can go straight to your agent.

You know what output you need
your agent, directly
The output, with the checks that ran against it.
Behavior or scope is unclear
plan · research
Plan settles the behavior. Research investigates sources and cites evidence.
An output or design already exists
review · validate
Review gives advice. Validate returns PASS, FAIL or NOT_PROVEN against acceptance.
Work spans tasks or agents
orchestrate
Disjoint scopes, one integrated output, checks and a fresh review where costly.
Earlier work may answer it
memory
Recall relevant context, or draft a report from recorded evidence.

Not sure where to start? Try the read-only Research task →

The reference implementation

AgentOps puts the factors on paved paths.

Its optional skills and the ao CLI work in the coding agent you already use. Start with a read-only Research task in a repository you know. This site tracks AgentOps v3.9.0.

v3.9.0 README ↗ · Release notes ↗

terminal · Claude Code
claude plugin marketplace add boshu2/agentops
claude plugin install agentops@agentops-marketplace

Codex and Cursor commands are on the install page.

Your agent said done. Something else proved it. The proof is yours.