At a glanceSecurity
Every defence sits before, around, or after the model — never inside the prompt. An attacker who reaches the model has already beaten anything you wrote there.
LandscapeTypes & approaches
Click a highlighted type to open its own page — concept, use case, and diagram.
The attacker gets a vote
Guardrails keep a cooperative system inside its lane. Security is what you need when someone is actively trying to push it out. Different threat model, different defences, and almost none of the intuition transfers.
Classical software has a boundary you can point at. Code is code, data is data, and the CPU never confuses one for the other — that separation is enforced by the machine, and every injection bug in the history of computing is a story about somewhere that separation quietly broke down (SQL injection, XSS, buffer overflows). We spent fifty years learning to defend that boundary.
An LLM has no such boundary. The system prompt, the user's message, the retrieved document, the tool result, and the web page the agent just fetched all arrive as one flat sequence of tokens. There is no bit that marks "this part is instructions" and "this part is only data to be read". The model infers the difference from context, statistically, and any statistical inference can be pushed. Greshake et al. put it exactly: LLM-integrated applications "blur the line between data and instructions" [3]. That single sentence generates the entire attack surface below.
This matters more the more capable your system is. A chatbot that can only talk has a reputational problem. A RAG system has a data problem. An agent with tools, credentials and network access has a breach problem — because the model is the thing deciding which tools to call, and the model is the thing the attacker just took over.
The honest baseline: OWASP states plainly that "it is unclear if there are fool-proof methods of prevention for prompt injection" [2]. Design for containment, not for a fix.
The mapOWASP Top 10 for LLM Applications (2025)
OWASP's GenAI Security Project maintains the reference list of LLM risks, refreshed for 2025 [1]. It is the closest thing the field has to a shared vocabulary, and it is what an auditor or an enterprise security review will ask you about by ID. Learn the IDs. The right-hand column is where each one is handled on this site.
| ID | Risk | What it means in your system | Where it bites hardest |
|---|---|---|---|
| LLM01 | Prompt Injection | Untrusted text alters model behaviour. Direct (the user types it) or indirect (it arrives inside retrieved content). | Anything that reads text you did not write |
| LLM02 | Sensitive Information Disclosure | The model emits PII, secrets, or another tenant's data — from context, from training, or from a tool result. | RAG over mixed-permission corpora |
| LLM03 | Supply Chain | Compromised model weights, datasets, adapters, or packages. | Every from_pretrained() call |
| LLM04 | Data and Model Poisoning | Adversarial content in pre-training, fine-tuning, or embedding data. | Any corpus users can write to |
| LLM05 | Improper Output Handling | Model output is passed unvalidated into a shell, SQL, browser, or renderer. | Agents, code interpreters, Markdown UIs |
| LLM06 | Excessive Agency | Too many tools, too much permission, too little confirmation. | The confused-deputy section below |
| LLM07 | System Prompt Leakage | Your instructions — and anything foolishly embedded in them — come back out. | Prompts holding keys, rules, or schemas |
| LLM08 | Vector and Embedding Weaknesses | Attacks on the retrieval layer itself: poisoned chunks, cross-tenant leakage, inversion. | Shared vector stores |
| LLM09 | Misinformation | Confident wrong output that a downstream system trusts. | Automated decisioning |
| LLM10 | Unbounded Consumption | Cost and availability attacks — token floods, agent loops, model theft by distillation. | See Costing |
Source: OWASP Top 10 for LLM Applications, 2025 revision [1]. The 2025 list added System Prompt Leakage and Vector/Embedding Weaknesses and renamed several 2023 entries — cite the year when you reference it, because the IDs moved.
LLM01 · the root bugPrompt injection: direct and indirect
Direct injection is the one everybody pictures: a user types something that overrides your system prompt. It is real, and it is the less interesting half. The user attacking their own session is mostly attacking their own privileges — annoying, cheap, usually contained.
Indirect injection is the one that should change your architecture. The malicious instruction is not typed by the user at all. It is already sitting in content your system retrieves and drops into the context window: a web page, a PDF, a support ticket, a calendar invite, a code comment, a tool response, an email. Greshake et al. demonstrated this against real systems — including Bing's GPT-4 chat and code-completion engines — and showed it enables remote exploitation with no attacker access to the interface at all [3]. The attacker does not need to talk to your app. They need to write something your app will read.
Sit with the consequence. Your users are not the threat surface. Your corpus is the threat surface. Any document that anyone outside your trust boundary can influence is an instruction channel into your model. If your RAG index ingests public web pages, your model takes instructions from the internet. If it ingests inbound support email, your model takes instructions from anyone who knows your support address.
Why the payload does not look like a payload
- Invisible to humans, legible to models.HTML comments, white-on-white text, zero-size fonts, alt-text, document metadata, and off-screen elements are all stripped from the human view and preserved in the text extraction that feeds your chunker.
- It reads like polite prose.The EchoLeak analysis notes the effective payload avoided obvious injection syntax and instead read as an ordinary business request [20]. Signature matching for "ignore previous instructions" catches a thing nobody serious does any more.
- It covers its tracks.Payloads routinely include an instruction not to mention the payload — so the summary the user reads looks clean [20].
- It survives your pipeline.Chunking, embedding and re-ranking do not sanitise anything. They faithfully carry the instruction into the prompt.
And it persists. An injected instruction that lands in conversation memory, a summary, a scratchpad file, or a re-indexed note becomes a stored injection — the LLM equivalent of stored XSS, firing on every future session that reads that record.
LLM04 + LLM08RAG poisoning: five documents against millions
If you build on this site's material you will build RAG. So this is the attack to internalise.
PoisonedRAG (USENIX Security 2025) is the load-bearing result. The attacker injects a small number of crafted texts into a knowledge base to force a chosen answer for a chosen question. The number required is the part that should stop you: a 90% attack success rate from injecting five malicious texts per target question, into a database containing millions of texts [4].
Five in millions. Your corpus size is not a defence. Dilution is not a defence. Retrieval is a search for the most relevant chunk, and the attacker gets to write the most relevant chunk.
That is the mechanism, and it is worth being precise about it because it inverts the intuition everyone brings. Retrieval is adversarially helpful: it is optimised to surface the passage that best matches the query. An attacker who knows the target question can craft a chunk that is a near-perfect semantic match and therefore gets promoted to the top of the results, past thousands of legitimate documents. The retriever is doing its job correctly. That is the problem.
The PoisonedRAG authors also evaluated defences and found them insufficient [4] — a finding that rhymes with everything else in this section.
Two distinct failure modes, often confused
- Knowledge corruption (LLM04).The poisoned chunk contains false facts. The model reads them, believes them, and answers wrongly with full confidence and a citation. No instruction is injected; nothing "breaks". Your evals pass. This is the quiet one.
- Indirect injection via retrieval (LLM01).The poisoned chunk contains instructions. The model reads them and acts — calls a tool, leaks context, rewrites its own task. This is the loud one, and the one that becomes EchoLeak.
Defending the retrieval layer
- Provenance on every chunk.Store source, author, ingestion time and trust tier as metadata at index time. You cannot make trust decisions later if you did not record trust earlier. This is the single highest-leverage thing on this page and it costs one column.
- Tier your corpus, and let tier drive privilege.First-party curated content is not the same as scraped web content is not the same as user-submitted content. Retrieval from a low-trust tier must downgrade what the agent is subsequently allowed to do — see Context-Minimization and Plan-Then-Execute below.
- Partition by tenant and permission at the index, not the prompt.A shared vector store with post-hoc filtering is a cross-tenant leak waiting for a bug (LLM08). Filter in the query, enforce in the store, and re-check the ACL of every retrieved chunk against the calling user before it enters the prompt.
- Gate the write path.Ask who can put a document into this index, and by what route. Most teams can answer for the manual path and not for the automated one — the crawler, the email connector, the Slack sync, the "helpful" integration someone shipped last quarter.
- Diversity checks over top-k.If one retrieved chunk contradicts the other k−1, that is signal, not noise. Cross-check high-stakes answers against multiple independent sources rather than the single best match.
- Re-scan on re-index.Poisoned content that entered six months ago is still in there. Scanning at ingestion only protects you going forward.
Alignment under attackJailbreaks: why the model itself cannot be the last line
A jailbreak targets the model's safety training rather than your application logic. For a builder, the point of studying jailbreaks is not the harmful-content angle — that is Guardrails territory. The point is what jailbreaks prove about model-level defences in general, which is what your injection defences also are.
Two results are worth knowing exactly.
Automated and transferable (GCG, 2023)
- Zou et al. used greedy and gradient-based search to automatically generate adversarial suffixes — no human creativity required, and therefore no bottleneck on attack volume [5].
- The finding that matters: suffixes optimised against open models (Vicuna-7B/13B) transferred to models the attacker never touched — ChatGPT, Bard, Claude, LLaMA-2-Chat, Pythia, Falcon [5].
- Consequence for you: closed weights are not a defence. An attacker can develop against a local open model, offline and unrate-limited, and carry the result to your production endpoint. Your API logs will show the first attempt, not the thousand rehearsals.
Long context is attack surface (Many-shot jailbreaking, 2024)
- Fill the context with hundreds of faux dialogues demonstrating the unwanted behaviour, and the model's in-context learning does the rest. Effectiveness follows a power law in the number of shots, tested up to 256 [6].
- The capability is the vulnerability: models that are better at in-context learning are more susceptible, and the attack only became practical because context windows grew from ~4,000 tokens to 1,000,000+ [6]. Every context-window upgrade you celebrate is also an attack-surface upgrade.
- The defence lesson is the valuable part. Fine-tuning the model to refuse these prompts only delayed the jailbreak — it took more shots, then it worked anyway. What worked was classifying and modifying the prompt before it reached the model, cutting attack success from 61% to 2% [6].
Generalise that: a defence inside the model competes with the model's own capabilities and tends to lose; a defence outside the model does not have to. That principle drives the architecture section below.
Model-layer defences are still worth buying — they are just not sufficient. Anthropic's Constitutional Classifiers, an external filtering layer, cut universal-jailbreak success from 86% to 4.4%; across ~3,000 hours of red-teaming by 185 participants, no participant found a universal jailbreak covering all ten target queries. The first generation cost 23.7% additional compute [7]. That is the real shape of the trade: a large, measurable risk reduction, a real bill, and not zero.
LLM02 + LLM05Data exfiltration through tool use
Injection alone changes what the model says. Injection plus an outbound channel changes what the attacker has. The channel is rarely a tool called send_data_to_attacker. It is almost always something innocuous.
CVE-2025-32711 ("EchoLeak"), Microsoft 365 Copilot, 2025 — described in the published analysis as the first known case of prompt injection weaponised for concrete data exfiltration in a production AI system, and genuinely zero-click [20]. Read the chain as a defender, because every link is a control that exists in your system too:
Delivery. An email arrives. The victim never opens it. It sits in the mailbox as ordinary retrievable content — and Copilot's RAG layer indexes the mailbox. The attacker's only requirement was to know an email address.
Classifier bypass. Microsoft ran XPIA, a dedicated cross-prompt-injection classifier. The payload read like a normal business request rather than an injection, and passed [20]. Detection was present and was not enough.
Trigger. The user later asks Copilot something ordinary. Retrieval pulls the attacker's email into context alongside genuinely sensitive internal data. The user did nothing wrong.
Output-filter bypass. Copilot redacted standard Markdown links [text](url). The payload used reference-style Markdown instead — a syntax variant the sanitiser did not recognise [20]. The allowlist was written against a format, not a capability.
Automatic egress. No click needed: a reference-style image tag means the browser fetches the URL — with the stolen data in the query string — the moment the answer renders [20]. Rendering is a network call.
CSP bypass. Content Security Policy blocked attacker domains, so the payload proxied through an allowed Microsoft service — a Teams URL-preview API — which dutifully fetched the data on Copilot's behalf [20]. Your allowlist inherits every open redirect and proxy inside it.
Six controls. Each individually reasonable. The chain went through all of them, and the published analysis concludes that only a layered, defence-in-depth approach can contain this class of threat, because the underlying issue is the collapse of the boundary between content and commands [20].
Exfiltration channels people forget to close
- Rendered Markdown images and linksthe classic. Any
in model output is an unauthenticated GET the instant it renders. - Allowlisted domains with proxies, previews, or open redirectsas above. Allowlist the URL, not the host, wherever you can.
- Tools with a URL parameter
fetch_url, webhooks, "share to", URL-shorteners, link-preview generators, image generation from a remote reference. - Legitimate write toolssending an email, filing a ticket, committing a file, posting to a channel the attacker can read. Data leaves through the front door.
- DNS and error messagesa hostname is a channel; a verbose stack trace is a channel.
- Slow channelsone secret token per session across many sessions still adds up, and no single request looks anomalous.
The lethal trifecta [21] is the compression of all of this into something you can hold in your head and say out loud in a design review. An agent is exposed when it combines: (1) access to private data, (2) exposure to untrusted content, and (3) the ability to communicate externally. All three together, and an attacker can trick it into reading your data and sending it out. Remove any one, and the attack loses its shape — which is a design instruction, not an observation. Most agents acquire all three by accident, one useful feature at a time, and nobody notices because no single pull request adds more than one.
LLM06 · the structural flawThe confused deputy in agent design
This one has a name and a literature because we solved it before — in 1988. Norm Hardy described the confused deputy: a program tricked by a less-privileged party into misusing its own authority. His example was a compiler with permission to write to a system directory; a user who asked it to write debug output to the billing file got the billing file destroyed, because the compiler used its permissions, not the user's. The subtitle of that paper is the punchline: "or why capabilities might have been invented" [22].
Your agent is the compiler. It holds credentials — a database connection, an API key, an OAuth token, a service account. It receives instructions from a channel an attacker can write to (see: your corpus). It acts with its own ambient authority. That is the 1988 bug, reissued with a chat interface, and it is why "the agent has admin so it can help anyone" is a sentence that should stop a design review.
Hardy's diagnosis transfers exactly: the vulnerability comes from ambient authority — permission that applies implicitly whenever the deputy acts, decoupled from the request that triggered it. The fix transfers too. Capability systems bundle the designation of an object together with the permission to access it [22], so authority travels with the request instead of hanging in the air.
What this means concretely
- Act as the user, not as the agent.Pass the caller's identity through to every tool and enforce authorisation at the resource, on the caller's rights. If your agent can read a record the requesting user cannot, you have built a privilege-escalation oracle with a friendly tone.
- Scope tokens per task, not per deployment.Short TTL, narrow scope, one purpose. The long-lived god-token in the environment variable is the whole vulnerability.
- Least privilege on the tool list.Every tool is permanent attack surface. The read-only variant of a tool is a different tool — ship that one.
- Enforce outside the model.The model asking "am I allowed?" is not authorisation, it is a suggestion box. Authorisation happens in code the model cannot influence.
- Human approval on consequential, irreversible actionsone of OWASP's explicit LLM01 mitigations [2]. Make the confirmation show the resolved action (this file, this recipient, this amount), never the model's summary of it. A model under injection writes a reassuring summary.
- Log the decision, not just the outcome.When the incident happens you need to know which retrieved chunk was in context when the tool fired. Log retrieved doc IDs with every tool call, or you will be unable to answer the only question that matters.
LLM03 + LLM04Supply chain: models, datasets, packages
Software supply chain security assumed artifacts are inert until executed. A model file is not inert. Serialised model formats can execute code on load, before any human or eval sees a single token of output.
Weights are code
- JFrog found around 100 malicious models on Hugging Face carrying genuine payloads (Feb 2024). The mechanism: PyTorch's pickle serialisation, abused via the
__reduce__method to run arbitrary code during deserialisation. Payloads opened a reverse shell to an external server — the model loads normally, nothing looks wrong, and the attacker has your machine [17]. - Defence: prefer
safetensors, which cannot carry executable payloads, and treat any pickle-based artifact from a source you do not control as untrusted code you are about to run as your service account. Because it is. - Do not over-trust scanners here. Pickle-scanning tooling has itself had bypass vulnerabilities. Format choice beats detection.
Names are not identities
- Model namespace reuse(Unit 42, 2025): when a Hugging Face author or org is deleted, the
Author/ModelNamenamespace can be re-registered by anyone. Every deployment that pulls that name by string then pulls the attacker's model instead. The researchers demonstrated it end-to-end, gaining a reverse shell on infrastructure hosting the model after a major cloud platform deployed the re-registered name [18]. - Disclosed to Google, Microsoft and Hugging Face; Google now scans daily for orphaned models. The core issue remains for anyone pulling models by name alone [18] — which is the default in every tutorial ever written.
- Defence:pin by commit hash, not name or tag. Verify checksums. Mirror what you depend on into storage you control. Treat a model reference exactly like a dependency reference, because it is one.
Training data is a trust boundary
- Anthropic, the UK AI Security Institute and the Alan Turing Institute trained 72 models (600M / 2B / 7B / 13B params, Chinchilla-optimal data) and found that ~250 poisoned documents sufficed to install a backdoor at every size tested — near-constant, not a percentage of the dataset [16].
- Read that carefully: a 13B model sees 20× the data of a 600M model, and needed the same ~250 documents. Scaling up your training set does not dilute the attack. The intuition that big data drowns bad data is simply wrong.
- Cite the caveats or you will be citing it wrong.The authors are explicit: this is a narrow denial-of-service backdoor (gibberish after a
<SUDO>trigger), unlikely to matter much in advanced models; whether the dynamics hold for more complex behaviours like backdooring code or bypassing guardrails, and for larger models, is unclear; and it concerns pretraining-scale poisoning, not post-training defences [16]. It is a strong reason to take fine-tuning-data provenance seriously. It is not evidence that frontier models are backdoored. - Defence:the exposure you actually control is your fine-tuning set. Know where every example came from. Be extremely careful with user-generated content, scraped data, and synthetic data generated by a model that itself read untrusted input.
The rest of the chain
- Packages.The ML ecosystem installs fast and pins loosely. Lockfiles, hash-pinning and a private index are table stakes, and typosquatting on ML package names is a live technique.
- Adapters and LoRAs.A LoRA is a behaviour patch from a stranger, applied to a model you trust. It inherits none of the base model's safety evaluation.
- Tool and connector ecosystems.A third-party tool definition is a set of instructions your model will read and follow, delivered by someone else, and mutable after you reviewed it. Pin versions and re-review on change.
- Keep an inventory.Models, datasets, adapters, and tools — with versions and hashes. You cannot patch what you cannot enumerate.
LLM02 + LLM07PII leakage and memorisation
Sensitive data reaches a user through three distinct doors. They have three distinct fixes and teams routinely fix one and claim all three.
| Door | Mechanism | Where the fix lives |
|---|---|---|
| From context | You retrieved data this user may not see, or another tenant's chunk, and the model repeated it. | Retrieval-time authorisation. Never a prompt instruction. |
| From training | PII was in the fine-tuning set and the model memorised it. | Data minimisation, redaction before training, DP techniques where warranted. |
| From the prompt | Your system prompt leaks (LLM07) — including anything unwise embedded in it. | Never put secrets in a prompt. Assume it is public. |
On memorisation, the strongest published result is Nasr et al.'s divergence attack: a technique that pushes a production model off its chatbot-style generation and causes it to emit training data at ~150× the normal rate. The authors extracted gigabytes of training data across open, semi-open and closed models, including ChatGPT [19]. Memorisation is not folklore, it is measured, and "the model would never say that" is not an argument.
Treat the system prompt as public. Not "probably fine" — public. LLM07 exists as its own OWASP entry because teams keep putting API keys, internal schemas, and business rules in prompts, and those prompts keep coming back out. If leaking your prompt would be a breach, your architecture is the breach.
Practical controls
- Minimise before you prompt.Redact or tokenise PII on the way in. Data that never enters the context cannot leave it.
- Authorise at retrieval, per request, per user.Re-check the ACL of every chunk against the calling user at query time. Index-time permissions go stale the moment someone changes teams.
- Scan the output toobut understand what that buys: it catches accidents and shapes you predicted. It is a backstop, not a boundary.
- Mind the logs.Prompt and completion logs, traces and eval datasets are copies of your most sensitive data in systems with weaker access control than the source. This is a very common real-world leak and it has nothing to do with the model.
- Check retention with your providerwhat is stored, for how long, whether it trains anything, and which region it sits in.
The honest tableWhat works, what only looks like it works
This is the section to read twice. The AI security market sells a lot of things that demonstrate beautifully and fail against an attacker who knows they are there.
Start with the one everybody tries first. "Just tell the model to ignore instructions in retrieved content." It is free, it is one line, it demos perfectly, and it does not work. Three reasons, and they compound:
Why instructing the model to ignore injections fails
- It is the same channel.Your defensive instruction and the attacker's instruction are both text in the same context window, competing for the same attention. You are not enforcing a rule; you are entering an argument, on equal terms, in a medium the attacker also writes fluently.
- Position and phrasing are attacker-controllable.The injection can arrive later in the context, appear more specific, claim more authority, impersonate a system message, or simply restate your rule and then declare an exception to it.
- It is untestable in the direction that matters.Passing your test suite tells you it stops the attacks you thought of. The attacker's job is the attacks you did not think of, and the search space is unbounded.
The empirical version of that argument is the most important defensive finding of 2025. Google DeepMind evaluated in-context defences (spotlighting, paraphrasing, in-context learning, self-reflection, perplexity filtering) and classifier defences against adaptive attackers — attackers who optimise against the deployed defence rather than against an undefended baseline. Their finding, verbatim: "Many defenses that perform well on our static evaluation set can be tricked by small and subtle adaptations to an attack." Their conclusion: "Robustness requires defense in depth. Adversarially training improves resilience to known attacks, however enumerating the space of all possible attacks is intractable" — and defences must sit at every layer of the stack [15].
Any vendor benchmark that does not describe an adaptive attacker is measuring the wrong thing. "Blocks 99% of prompt injections" against a static corpus is a claim about a corpus, not about your security.
| Defence | What it actually buys | How it fails | Verdict |
|---|---|---|---|
| "Ignore injected instructions" in the system prompt | Nothing an attacker respects. Slightly fewer accidents. | Same channel, same tokens, attacker writes last and louder. | Theatre. Costs nothing, so keep it, but never count it. |
| Keyword / regex filters ("ignore previous instructions") | Blocks copy-pasted 2023 payloads and script kiddies. | Paraphrase, encode, translate, split across chunks. EchoLeak's payload read as normal business prose. | Theatre against a real attacker. |
| Delimiters alone ("data is between <doc> tags") | Marginal clarity for the model. | The attacker writes </doc>. It is the quoting bug you already know from SQL. | Necessary, nowhere near sufficient. |
| Spotlighting (delimiting, datamarking, encoding) | Real, measured reduction: attack success from >50% to <2% with minimal task-quality impact [9]. | An in-context defence — degrades under adaptive attack [15]. | Worth deploying. A layer, not a boundary. |
| Injection classifiers (input and output) | Genuine risk reduction; catches volume. Constitutional Classifiers: 86% → 4.4% on universal jailbreaks, at 23.7% compute [7]. Prompt classification cut many-shot success 61% → 2% [6]. | Microsoft ran a dedicated classifier (XPIA) and EchoLeak walked past it [20]. Adaptive attackers optimise against the classifier. | Buy it. Do not bet on it. |
| Self-reflection ("check if you were manipulated") | Catches clumsy attacks. | The compromised model is the auditor. Evaluated by DeepMind among the defences vulnerable to adaptation [15]. | Weak. Never load-bearing. |
| Instruction hierarchy (trained-in privilege levels) | Trains the model to prioritise privileged instructions and selectively ignore lower-privileged ones; "drastically increases robustness — even for attack types not seen during training" [8]. | Still probabilistic. Improves the odds; proves nothing. | Use models that have it. Not a boundary. |
| StruQ / SecAlign (structured queries + preference optimisation) | The strongest model-level numbers published. Both drive a dozen-plus optimisation-free attacks to ~0%. Against optimisation-based attacks StruQ sits at 45% ASR; SecAlign cuts it to 8%, preserving AlpacaEval2 utility on Llama3-8B-Instruct [10][11]. | 8% is not 0%. Requires fine-tuning; you need control of the weights. | Best-in-class model layer. Still a layer. |
| Least privilege + user-scoped tokens | Bounds the blast radius regardless of whether injection succeeds. | Fails only if you granted too much — which is a decision you can audit, not a probability. | Works. Do this first. |
| Breaking the lethal trifecta | Removes the attack's shape: no private data, or no untrusted content, or no egress [21]. | Requires saying no to a feature. That is the actual difficulty. | Works. The highest-leverage design decision on this page. |
| Deterministic egress control (URL allowlist, no auto-fetch) | Enforced in code the model cannot argue with. | Only as good as the allowlist — proxies, previews and open redirects inside it are holes [20]. | Works, if you audit the allowlist. |
| Human approval on consequential actions | A real boundary — OWASP names it explicitly [2]. | Approval fatigue; and if you show the model's summary rather than the resolved action, you are asking the human to approve the attacker's prose. | Works if rare and specific. |
| Architectural isolation (CaMeL, Dual-LLM, Plan-Then-Execute) | Security properties from structure rather than from model behaviour [12][13]. | Costs capability and engineering effort. Not every task fits the pattern. | The only approach with provable properties. |
Numbers are from the cited papers under their own evaluation conditions and do not transfer to your system unchanged. Read a defence's threat model before you read its percentage — and check whether the attacker was allowed to adapt.
The way outDesign patterns with provable properties
Every model-layer defence above is probabilistic: it lowers the odds. The research consensus is that you get real guarantees only by changing the architecture so that a successful injection cannot reach anything that matters.
The organising principle, from Beurer-Kellner et al. — and this is the sentence to remember from this entire page:
"Once an LLM agent has ingested untrusted input, it must be constrained so that it is impossible for that input to trigger any consequential actions." [13]
Note the word impossible, not unlikely. The paper starts from the assumption that the model is already compromised — an agent may do anything after ingesting untrusted input — and asks what structure survives that assumption. Their six patterns, each trading some generality for containment [13]:
| Pattern | The constraint | Use when |
|---|---|---|
| Action-Selector | The LLM may only pick from a prescribed list of actions and arguments. It cannot emit free text back to the caller, so injected content has no channel to act through. | Routing, triage, classification — the action space is genuinely enumerable. |
| Plan-Then-Execute | The agent commits to a fixed plan before touching untrusted data. Tool results can change arguments; they cannot add steps. | Most real workflows. The best generality-to-safety ratio here. |
| LLM Map-Reduce | An isolated LLM processes each untrusted document separately; only constrained, structured outputs are aggregated. One poisoned document cannot reach across. | Summarising or extracting across a corpus you do not fully trust. |
| Dual LLM | A privileged LLM never sees untrusted content. A quarantined LLM reads it and returns only symbolic references the privileged one handles blind. | Agents that must combine private data with untrusted input. |
| Code-Then-Execute | The LLM writes a program describing the whole task; a deterministic interpreter runs it and enforces the data-flow rules. This is CaMeL's shape. | Complex multi-step tasks where you need real guarantees. |
| Context-Minimization | Strip untrusted content out of the context once it has served its purpose, rather than containing its effects afterwards. | Cheap and composable. Combine with anything above. |
CaMeL is the most complete instantiation and the one to study. It builds a protective system layer around the LLM that holds "even when underlying models are susceptible to attacks". It extracts control and data flow from the trusted query, so untrusted data can influence values but never the program's structure, and uses capabilities — the direct descendant of Hardy's 1988 answer — to block unauthorised exfiltration through tool calls [12].
And it publishes its price honestly, which is why it is worth citing in a design review: on AgentDojo, CaMeL solved 77% of tasks with provable security, versus 84% for an undefended system [12]. Seven points of capability for a security property you can actually reason about. That is the trade, stated plainly. Most vendors will not state theirs.
Measure your own system: AgentDojo is the open benchmark for exactly this — 97 realistic tasks across email, banking and travel, with 629 security test cases, built as a dynamic environment so you can add adaptive attacks rather than replay a fixed corpus [14]. Its baseline finding is a useful humility check: current LLMs fail many of these tasks even with no attacker present.
In practiceWhat to actually do
Ordered by leverage. If you do the first three and nothing else, you have eliminated most of what this page describes.
Draw the trifecta on your own architecture. Mark every component that touches private data, every path untrusted content can enter, and every route out — including image rendering, link previews, DNS, and any allowlisted host with a proxy. Where all three meet, you have a decision to make, and it is a product decision, not a security one [21].
Kill ambient authority. The agent acts as the user, never as itself. Scope tokens per task, short TTL. Enforce authorisation at the resource, in code, on the caller's rights — never by asking the model [22].
Pick a pattern before you pick a model. Plan-Then-Execute or Action-Selector covers most real workflows and gives you containment that does not depend on the model being un-fooled [13].
Tag provenance at ingestion. Source, trust tier, timestamp on every chunk — and let tier drive privilege downstream. You cannot retrofit this; a corpus without provenance metadata has to be re-indexed to get it [4].
Make egress deterministic. URL allowlist enforced in code. Audit the allowlist for proxies, previews and open redirects. Disable auto-fetch of remote images in rendered output. Strip or neutralise Markdown link and image syntax in model output — all the syntaxes, including reference-style [20].
Pin the supply chain. Models by commit hash, never by name. safetensors over pickle. Lockfiles and hash-pinned packages. Mirror what you depend on. Keep an inventory of models, datasets, adapters and tools with versions [17][18].
Layer the probabilistic defences anyway. Spotlighting, injection classifiers, instruction-hierarchy-trained models, SecAlign-style training if you own the weights. Each cuts volume substantially. None is a boundary. Deploy them behind the architectural controls, not instead of them [9][10][15].
Test adaptively, and keep testing. Run AgentDojo or an equivalent in CI. Red-team with an attacker who knows your defences. A static attack corpus measures yesterday [14][15].
Log for the incident you will have. Every tool call, with the retrieved document IDs that were in context when it fired. Without that link you cannot answer "which document made it do that", which is the first question and often the only one that matters.
Assume the system prompt is public and the logs are sensitive. No secrets in prompts. Prompt/completion logs, traces and eval sets get the same access controls as the source data they copy [19].
The one-line version: you cannot make the model un-foolable, so make being fooled not matter. Every durable defence on this page is a variation of that sentence.
Questions you will be askedStraight answers
Can we just solve prompt injection with a better system prompt?
No. Your instruction and the attacker's instruction are the same kind of object in the same context window. OWASP: "it is unclear if there are fool-proof methods of prevention" [2]. Constrain what a fooled model can reach.
Our corpus is internal-only. Are we safe from RAG poisoning?
Ask who can write into it. Internal wikis take pasted content from the web. Ticket systems take inbound email. Repos take dependency READMEs. Nearly every "internal" corpus has an external write path someone forgot about — and PoisonedRAG needed five documents [4].
We use a closed frontier model. Doesn't that stop adversarial attacks?
No. GCG suffixes optimised against open models transferred to ChatGPT, Bard and Claude [5]. The attacker develops offline against open weights and arrives at your endpoint with a working attack on the first request.
We bought an AI firewall / injection classifier. Are we done?
You have bought a useful layer. Microsoft had a dedicated cross-prompt-injection classifier and EchoLeak went through it [20], and DeepMind found defences that look strong statically fall to small adaptations [15]. Ask the vendor for adaptive-attacker results. If they cannot produce them, you have a static-corpus number.
Isn't the 250-documents poisoning result terrifying?
It is important and it is narrower than the headlines. It is a gibberish-on-trigger backdoor at pretraining scale; the authors say explicitly that it is unlikely to matter much in advanced models, that complex behaviours are unproven, and that it says nothing about post-training defences [16]. What it should change is your handling of your own fine-tuning data — because "our dataset is huge" is now definitively not a defence.
What does security cost us?
Quantifiably. CaMeL: 77% vs 84% task success for provable security [12]. First-gen Constitutional Classifiers: 23.7% additional compute [7]. Present it as a trade with numbers, not as a tax, and note that the alternative is an unbounded liability with no number at all.
Where does this stop being Security and start being Guardrails?
Guardrails assume the system wants to comply and might drift. Security assumes someone is optimising against you. If the threat has a budget and a feedback loop, it is this page. See Guardrails for the safety and compliance half.
ReferencesSources
Every external claim on this page traces to one of these. Security moves fast — the papers below are dated on purpose. Re-check the OWASP list and the agent-defence literature at least twice a year; a defence that was state of the art in 2024 may be a known-broken one now.
Cited work
- OWASP Top 10 for LLM Applications 2025 — OWASP GenAI Security Project
- OWASP LLM01:2025 Prompt Injection
- Greshake et al. (2023), Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Zou, Geng, Wang, Jia (2024), PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models (USENIX Security 2025)
- Zou, Wang, Carlini, Nasr, Kolter, Fredrikson (2023), Universal and Transferable Adversarial Attacks on Aligned Language Models
- Anil et al. (2024), Many-shot Jailbreaking — Anthropic
- Anthropic (2025), Constitutional Classifiers: Defending against universal jailbreaks
- Wallace, Xiao, Leike, Weng, Heidecke, Beutel (2024), The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Hines, Lopez, Hall, Zarfati, Zunger, Kiciman (2024), Defending Against Indirect Prompt Injection Attacks With Spotlighting
- Chen, Zharmagambetov, Mahloujifar, Chaudhuri, Wagner, Guo (2025), SecAlign: Defending Against Prompt Injection with Preference Optimization (ACM CCS 2025)
- Berkeley BAIR (2025), Defending against Prompt Injection with Structured Queries (StruQ) and Preference Optimization (SecAlign)
- Debenedetti, Shumailov, Fan, Hayes, Carlini, Fabian, Kern, Shi, Terzis, Tramer (2025), Defeating Prompt Injections by Design (CaMeL)
- Beurer-Kellner et al. (2025), Design Patterns for Securing LLM Agents against Prompt Injections
- Debenedetti, Zhang, Balunovic, Beurer-Kellner, Fischer, Tramer (2024), AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Shi et al. (2025), Lessons from Defending Gemini Against Indirect Prompt Injections — Google DeepMind
- Anthropic, UK AI Security Institute, Alan Turing Institute (2025), A small number of samples can poison LLMs of any size
- JFrog (2024), Data Scientists Targeted by Malicious Hugging Face ML Models with Silent Backdoor
- Palo Alto Unit 42 (2025), Model Namespace Reuse: An AI Supply-Chain Attack Exploiting Model Name Trust
- Nasr et al. (2023), Scalable Extraction of Training Data from (Production) Language Models
- EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit in a Production LLM System (2025) — analysis of CVE-2025-32711, Microsoft 365 Copilot
- Simon Willison (2025), The lethal trifecta for AI agents: private data, untrusted content, and external communication
- Norm Hardy (1988), The Confused Deputy (or why capabilities might have been invented) — ACM SIGOPS OSR
What changedWhat changed here
Updated this page Platform terms, not just technical capability, can block an agent from completing a purchase on a user's behalf, so agent checkout flows need per-site permission rather than assumed open access.
Updated this page The argument for MCP is control and auditability rather than capability, which is a counterpoint to wiring agents directly to APIs.
Three kinds of claim, strongest first. Signal runs every morning.