Skip to content
BITBRIEF

Institutional research · AI · Cybersecurity · Digital assets

Vol. 01 · No. 13

Prompt injection

An attack that places instructions in content a language model reads, causing the model to follow the attacker's directions instead of the operator's.

A model given a web page, a document or an email to process cannot reliably separate the data it was asked to work on from instructions hidden inside that data. Text that says «ignore your previous instructions» is, to the model, just more text.

Why it does not have a clean fix

Conventional injection attacks are solved by separating code from data at a boundary the interpreter enforces. A language model has no such boundary: instructions and content arrive in the same channel and are processed by the same mechanism. Mitigations reduce the rate; none of them close the class.

The consequence is proportional to the agent's reach

A model that only writes text produces bad text. A model that can call tools, spend money or modify systems does the attacker's work with the operator's credentials. This is why limits belong outside the agent — a budget the model can reason about is a budget the model can be talked past.

All 30 terms