Skip to content
BITBRIEF

Institutional research · AI · Cybersecurity · Digital assets

Vol. 01 · No. 13

How to secure an AI agent that can spend money

A control design, not a list of best practices. The threat model is that the agent will be instructed by an attacker, and the question is what that buys them.

In short

Design on the assumption that the agent will, at some point, follow an attacker's instructions. Prompt injection has no clean fix, so a control that depends on the model resisting instruction is not a control.

The exposure of any agent equals what it can reach, not how good the model is. Bounding reach is the only measure that reliably changes the outcome.

Three controls do most of the work: a spending ceiling enforced outside the agent, credentials scoped to the task rather than to the agent, and a human on anything irreversible above a threshold.

Log what the agent read, not only what it did. Without the inputs, an incident cannot be reconstructed — and the reasoning trace is not evidence, because it can be manipulated by the same attack.

An agent that can spend is a system with an attack surface made of text. Every document it reads, every page it fetches, every tool response it receives is a channel through which instructions can arrive. The model cannot reliably tell those instructions from the data it was asked to process, and no published defence has closed that class.

So the design question is not how to stop the agent being instructed. It is what an attacker gains when they succeed.

The threat model, stated plainly

AssumptionConsequence for design
The agent will follow an attacker's instruction at some pointControls that rely on the model refusing are not controls
The attacker controls text the agent reads, not the agent's codeApplication-level enforcement still works; prompt-level does not
The reasoning trace can be shaped by the same attackDo not use it as evidence of intent
The attacker cannot forge the wallet balance or the credential scopeEnforcement belongs there
What to assume, and what follows

Control 1: the ceiling lives outside

Fund a wallet rather than connecting a line of credit, or issue a credential scoped so tightly that exceeding it fails at the network rather than at the model. The property that matters is that the enforcing component never reads prose. A wallet with nothing left in it cannot be persuaded.

Set the ceiling at an amount you would accept losing entirely in a single incident, not at the amount the agent is expected to need. Those two numbers are different, and the first is the one that bounds your exposure.

Control 2: scope credentials to the task

An agent that reads invoices does not need write access to the ledger. An agent that buys data lookups does not need the ability to open new vendor relationships. Scoping to the task rather than to the agent is the difference between an incident that costs a capped amount and one that costs whatever the agent's identity could reach.

  • One credential per task, not one per agent.
  • Expiry by default, renewed by the system rather than held indefinitely.
  • Read and write separated, so summarising cannot become modifying.
  • No standing access to anything the agent needs only occasionally.

Control 3: a human on the irreversible

Approval does not prevent injection; it stops the damage. That distinction is worth keeping straight, because the value of a human in the loop is entirely in the second half and nothing in the first.

Put the threshold where the review is still meaningful. An approval step that fires forty times a day becomes a reflex, and a reflex is not a control. Better to set it high enough that each one gets read.

Control 4: treat tool output as untrusted input

Agents trust their own tools, which is reasonable until a tool returns attacker-controlled text. Search results, fetched pages, file contents, API responses from third parties — all of these are content somebody else may have written, arriving through a channel the agent has no reason to doubt.

The practical rule: any tool that can return text originating outside your systems is an untrusted input boundary, and should be treated with the same suspicion as a user upload.

What to log

RecordWhy
The instruction the agent was givenDistinguishes a mandated action from an injected one
Every input the agent readWhere the injection will be found
Tool calls and their responsesThe most common delivery channel
Each payment with the instruction that caused itTies the ledger entry to a decision
The reasoning trace, marked as unreliableUseful context; not evidence
Recording enough to reconstruct an incident

Most agent logging captures actions and omits inputs. That produces a record showing exactly what happened and giving no way at all to establish why — which is the question an incident review exists to answer.

A design lesson from orchestration research

Where several agents are coordinated by a central planner built on a language model, the planner inherits the same weakness. Research posted in August 2026 on market-based alternatives measured it: in a centralised allocator, inserting a single preference nearly doubled the favoured agent's share of tasks.

One line in a context window moved the distribution of work that far. The argument for auction-based orchestration is usually made on efficiency; the security argument is stronger, because a bidding mechanism moves the routing decision into a rule that does not read prose.

The failure mode nobody plans for

When an incident does happen, investigating it means asking a model to reason about attack code — and hosted commercial models may refuse. Hugging Face hit exactly this in July 2026 and ended up running an open-weight model on its own hardware, because the attacker was bound by no usage policy while the defender's forensic work was blocked by guardrails.

Worth deciding in advance rather than during: what you will use for analysis if your usual model declines the question.

Questions

Is there any way to stop prompt injection outright?
No. A language model has no boundary between instructions and data — both arrive as tokens in one context and are processed by the same mechanism. Mitigations lower the rate; none close the class. Design for a successful injection and bound what it can do.
Should the agent have its own identity or share a service account?
Its own, and per task where practical. A shared service account makes attribution impossible after the fact and makes scoping impossible before it — you cannot limit an identity to one task if several tasks use it.
How much should the spending ceiling be?
An amount you would accept losing entirely in one incident. That is a different number from what the agent is expected to spend, and it is the one that actually bounds exposure. Start low and raise it on evidence rather than on projection.
Can we rely on the agent's own reasoning log to detect misuse?
Not as evidence. The same attack that redirects the agent's actions can shape the explanation it produces, so a plausible-looking rationale is not confirmation that the action was legitimate. Keep the trace for context and base detection on inputs and outcomes.

All guides