An agent that can spend is a system with an attack surface made of text. Every document it reads, every page it fetches, every tool response it receives is a channel through which instructions can arrive. The model cannot reliably tell those instructions from the data it was asked to process, and no published defence has closed that class.
So the design question is not how to stop the agent being instructed. It is what an attacker gains when they succeed.
The threat model, stated plainly
| Assumption | Consequence for design |
|---|---|
| The agent will follow an attacker's instruction at some point | Controls that rely on the model refusing are not controls |
| The attacker controls text the agent reads, not the agent's code | Application-level enforcement still works; prompt-level does not |
| The reasoning trace can be shaped by the same attack | Do not use it as evidence of intent |
| The attacker cannot forge the wallet balance or the credential scope | Enforcement belongs there |
Control 1: the ceiling lives outside
Fund a wallet rather than connecting a line of credit, or issue a credential scoped so tightly that exceeding it fails at the network rather than at the model. The property that matters is that the enforcing component never reads prose. A wallet with nothing left in it cannot be persuaded.
Set the ceiling at an amount you would accept losing entirely in a single incident, not at the amount the agent is expected to need. Those two numbers are different, and the first is the one that bounds your exposure.
Control 2: scope credentials to the task
An agent that reads invoices does not need write access to the ledger. An agent that buys data lookups does not need the ability to open new vendor relationships. Scoping to the task rather than to the agent is the difference between an incident that costs a capped amount and one that costs whatever the agent's identity could reach.
- One credential per task, not one per agent.
- Expiry by default, renewed by the system rather than held indefinitely.
- Read and write separated, so summarising cannot become modifying.
- No standing access to anything the agent needs only occasionally.
Control 3: a human on the irreversible
Approval does not prevent injection; it stops the damage. That distinction is worth keeping straight, because the value of a human in the loop is entirely in the second half and nothing in the first.
Put the threshold where the review is still meaningful. An approval step that fires forty times a day becomes a reflex, and a reflex is not a control. Better to set it high enough that each one gets read.
Control 4: treat tool output as untrusted input
Agents trust their own tools, which is reasonable until a tool returns attacker-controlled text. Search results, fetched pages, file contents, API responses from third parties — all of these are content somebody else may have written, arriving through a channel the agent has no reason to doubt.
The practical rule: any tool that can return text originating outside your systems is an untrusted input boundary, and should be treated with the same suspicion as a user upload.
What to log
| Record | Why |
|---|---|
| The instruction the agent was given | Distinguishes a mandated action from an injected one |
| Every input the agent read | Where the injection will be found |
| Tool calls and their responses | The most common delivery channel |
| Each payment with the instruction that caused it | Ties the ledger entry to a decision |
| The reasoning trace, marked as unreliable | Useful context; not evidence |
Most agent logging captures actions and omits inputs. That produces a record showing exactly what happened and giving no way at all to establish why — which is the question an incident review exists to answer.
A design lesson from orchestration research
Where several agents are coordinated by a central planner built on a language model, the planner inherits the same weakness. Research posted in August 2026 on market-based alternatives measured it: in a centralised allocator, inserting a single preference nearly doubled the favoured agent's share of tasks.
One line in a context window moved the distribution of work that far. The argument for auction-based orchestration is usually made on efficiency; the security argument is stronger, because a bidding mechanism moves the routing decision into a rule that does not read prose.
The failure mode nobody plans for
When an incident does happen, investigating it means asking a model to reason about attack code — and hosted commercial models may refuse. Hugging Face hit exactly this in July 2026 and ended up running an open-weight model on its own hardware, because the attacker was bound by no usage policy while the defender's forensic work was blocked by guardrails.
Worth deciding in advance rather than during: what you will use for analysis if your usual model declines the question.
Questions
- Is there any way to stop prompt injection outright?
- No. A language model has no boundary between instructions and data — both arrive as tokens in one context and are processed by the same mechanism. Mitigations lower the rate; none close the class. Design for a successful injection and bound what it can do.
- Should the agent have its own identity or share a service account?
- Its own, and per task where practical. A shared service account makes attribution impossible after the fact and makes scoping impossible before it — you cannot limit an identity to one task if several tasks use it.
- How much should the spending ceiling be?
- An amount you would accept losing entirely in one incident. That is a different number from what the agent is expected to spend, and it is the one that actually bounds exposure. Start low and raise it on evidence rather than on projection.
- Can we rely on the agent's own reasoning log to detect misuse?
- Not as evidence. The same attack that redirects the agent's actions can shape the explanation it produces, so a plausible-looking rationale is not confirmation that the action was legitimate. Keep the trace for context and base detection on inputs and outcomes.