Own the Verdict, Rent the Harness
Keep acceptance and evidence under caller control, whichever runtime, orchestrator or factory runs the work.
Outputs and memory must remain understandable when runtimes change. Keep acceptance under the caller's control, then enforce that boundary when Factor XI stops unsafe work.
The rule
Own the criteria and evidence by which work is accepted.
The harness is the software that runs and coordinates the agents: a coding agent's runtime, an orchestrator or a factory. Select one for the task, and replace it when a better one arrives, without handing it authority over those criteria.
A native agent and shell can be enough. A factory such as Gas City can coordinate larger workloads when the caller selects it. Scheduling, worker completion, agreement among agents and a vendor's benchmark score do not establish that an output meets intent. Keep source identities and check evidence accessible through the tools that own them.
The evidencepractice
No study tests this rule directly. Research does show that completion signals and benchmark scores overstate accepted work, and that one model scores differently in different harnesses.
- SkillsBench, 2026 ran the same models in more than one harness. With skills, Gemini 3.1 Pro passed 60.8% of tasks in Gemini CLI and 52.8% in OpenHands. Claude Opus 4.7 passed 61.2% in Claude Code and 53.1% in OpenHands.
- The Holistic Agent Leaderboard, 2025 ran 21,730 agent rollouts across 9 models and 9 benchmarks. Reviewing the logs, its authors found agents searching for the benchmark online instead of solving the task. Higher reasoning effort lowered accuracy in the majority of runs.
- Lee et al., 2025 surveyed 319 knowledge workers. Higher confidence in generative AI was associated with less critical thinking.
Put it to work
- Keep accepted intent and required checks independent of runtime branding.
- Use existing source and tracker interfaces rather than a second control plane.
- Operate a selected factory through its own coordinator and supported interfaces.
- Preserve an inspectable output and evidence when changing tools.
- Judge a new model or harness on your own tasks and criteria before relying on it.
Example · the monthly operations report
The report uses the same ledger, reconciliation rule and approval boundary whether one agent or a selected factory produces it. A factory's “run complete” message cannot replace those checks or authorize publication.
Failure signal
- Runtime success becomes the only evidence of completion.
- Changing harnesses loses the source links or acceptance criteria.
- An agent bypasses the factory coordinator to create competing workers.
Done looks like
The caller can judge the output using the same criteria and sources regardless of which authorized runtime produced it.