Cap the Blast Radius
Bound access, actions and spending with limits the runtime or source system can enforce.
A bounded job needs bounded access and spending. Those limits constrain the paved path selected in Factor V.
The rule
Limit what an agent can read, change, publish and spend.
Match permissions and checks to the consequences of a mistake. A read-only investigation, a document edit and a production action need different boundaries.
Use the runtime or source system to enforce consequential limits. A sentence in a prompt does not enforce a spending cap or prevent publication. New contexts, retries and extra workers do not renew the caller's limits.
Content an agent reads can carry someone else's instructions. Do not let one unsupervised agent combine untrusted input, access to private data or sensitive systems, and the ability to change state or send data out. Remove one of the three, or put a person on the consequential action.
The evidencemeasured
Security researchers have tested whether instructions and model-side defenses stop a determined attacker. The published defenses did not.
- Nasr et al., 2025 attacked 12 recent defenses against jailbreaks and prompt injection with adaptive methods. Most fell to attack success rates above 90%. Most had originally reported rates near zero.
- Meta, 2025 calls prompt injection an unsolved weakness of all LLMs and sets a rule of two. Within one session an agent gets at most two of these: untrustworthy input, access to sensitive systems or private data, and the ability to change state or communicate externally. An agent that needs all three needs supervision. Willison, 2025 describes the same combination as the lethal trifecta.
- OWASP, 2025 lists excessive agency among its top ten risks for LLM applications and traces it to excessive functionality, permissions or autonomy.
Enforcement outside the model holds up better. Debenedetti et al., 2025 built a system layer that stops untrusted data from changing an agent's control flow. It solved 77% of benchmark tasks with provable security, against 84% for the undefended system.
The spending, time and retry limits in this factor are operating practice.
Put it to work
- Grant access only to the sources and actions needed for the job.
- Separate preparation from consequential publication or execution.
- Set real time, cost or attempt limits when the task needs them.
- State the stopping condition and who can authorize a wider scope.
- Find the jobs that combine untrusted input, private data and outbound action; remove one or add approval.
Example · the monthly operations report
The report agent may read the approved ledger and write a private draft. It cannot email customers, change ledger entries or publish the report. A runtime spending limit bounds the investigation; reaching it produces an incomplete handoff.
Failure signal
- A drafting task receives unnecessary write or send permissions.
- Retries continue beyond the original limit.
- The agent treats access as permission to act.
- One agent reads untrusted content, holds private data and can send it out.
Done looks like
The job runs inside explicit, enforceable boundaries, and a stopped run states what remains unfinished.