Factor V of XIILay the Rails

Map the Terrain, Then Pave It

Inspect current sources, then encode supported methods in skills, templates and checks.

In the loop

The job's limits guide source discovery and method selection. The resulting paved path defines the resources that need one writer in Factor VI.

The rule

Understand the work before encoding a repeatable method.

Read current source contracts, tool behavior and relevant prior evidence. Use skills, templates and executable checks to make a supported practice easy to follow.

A paved path connects accepted intent to an inspectable output. It may serve research, reporting, operations or software delivery. Add a rule or gate for an observed failure with a concrete consumer; a process document alone does not improve the output. Keep guidance focused and specific, and have a person curate it. Where a rule must hold every time, make it an executable check.

The evidencemeasured

Benchmarks compare the same agents with and without written guidance.

  • SkillsBench, 2026 ran 87 tasks with and without curated skills across 18 model and harness combinations. Curated skills raised the average pass rate from 33.9% to 50.5%. Read that as a best case: the benchmark rejected tasks where skills made no measurable difference, and its authors call the set an optimistic scenario. Tasks given at most three skill modules gained more than tasks given larger bundles.
  • In the same study, skills the agents wrote for themselves scored below the no-skills baseline in all three configurations tested. Curated skills lowered the score on 13 of the 87 tasks.
  • Gloaguen et al., 2026 found repository context files had no significant effect on coding-agent success. Developer-written files significantly outperformed generated ones. The authors conclude that any change meant to improve performance should be evaluated before deployment.
  • The cost findings conflict. Gloaguen et al. measured more than 20% higher inference cost with context files. Lulla et al., 2026 found an AGENTS.md file associated with 28.64% lower median runtime at comparable task completion, across 124 pull requests.

Across these studies, guidance curated by a person beat guidance agents wrote for themselves, except where the repository had no other documentation. The case for shorter guidance is weaker. The largest bundles gained least, and that comparison is observational.

Put it to work

  • Inspect the existing workflow and source owners before inventing one.
  • Select an optional skill to resolve a concrete uncertainty.
  • Encode repeatable checks where they can catch a real defect.
  • Test reusable guidance on later work; revise or remove it when it does not help.

Example · the monthly operations report

Investigate why last month's report omitted reversals. Confirm the ledger schema, then add a reconciliation check and a report instruction covering reversals. The next report demonstrates whether that method catches the omission.

Failure signal

  • A template assumes fields the source does not contain.
  • A skill repeats advice but cannot guide a concrete decision.
  • A new gate adds ceremony without detecting the observed defect.
  • An instruction file grows with each run and nobody has measured whether it helps.

Done looks like

The method follows current sources, stays within authority and makes the required output easier to produce and check.