Prompt-injection defense stops text the model reads, not just text the user types, from taking over its instructions.
ConceptWhat it is
Prompt injection is an attack where malicious instructions are hidden inside user input or retrieved content, such as a document, webpage, or email the model reads, tricking it into ignoring its system prompt and following the attacker's instructions instead. It exists as a threat because LLMs cannot reliably distinguish trusted developer instructions from untrusted data in the same context window.
Defense combines input sanitization, instruction hierarchy, and output validation, since no single technique fully closes the gap.
How it worksThe mechanics
Systems layer defenses: delimiting untrusted content clearly, using models trained to respect an instruction hierarchy so system prompts outrank injected text, scanning inputs for known injection patterns, and validating that outputs and tool calls stay within an allowed action set before execution.
At a glanceSee it
Prompt injection splits into direct delivery and indirect delivery hidden in retrieved content — the zero-click indirect path is the dangerous one, since the victim never typed anything hostile, yet whatever the goal, the payoff usually rides one exit: an allowed action turned against you.
The dual-LLM pattern isolates untrusted text in a tool-less reader and lets a privileged planner act only on structured values — injected instructions never reach the tool boundary, unless raw content leaks back into the planner's own prompt.
When to use itWhere it fits
- Agentic systems that read untrusted web pages, emails, or documents and take actions.
- RAG applications where retrieved chunks come from external or user-uploaded sources.
- Any assistant with tool-calling access to sensitive actions like sending emails or executing code.
- Browser-using agents that navigate and interact with arbitrary websites.
When NOT to use itLimits & anti-patterns
- Closed systems that only process fully trusted, internally authored input, where the attack surface barely exists.
- Simple Q&A bots with no tool access, where a successful injection has no dangerous action to trigger.
Trade-offsAdvantages & costs
Advantages
- Reduces the risk of an agent being hijacked into leaking data or taking unauthorized actions.
- Instruction-hierarchy training is now built into many frontier models by default.
- Layered defenses catch different attack variants that a single filter would miss.
Trade-offs & costs
- No known technique provides complete protection against novel injection phrasing.
- Aggressive sanitization can break legitimate use cases that need to quote or process suspicious-looking text.
- Requires ongoing red-teaming as attackers develop new bypass techniques.
ExampleIn the real world
Anthropic's and OpenAI's browser-agent products restrict what an agent can do after visiting a webpage, requiring explicit user confirmation before actions like purchases or account changes, precisely to blunt indirect prompt injection from malicious page content.
ToolsHow to implement it
- Rebuffopen-source prompt-injection detection with canary tokens.
- Lakera Guardcommercial real-time injection and jailbreak detection API.
- Meta's Prompt Guarda lightweight classifier model for detecting injection and jailbreak attempts.
- NVIDIA NeMo Guardrailsa framework for enforcing conversational rails and action boundaries.
Cost & effortWhat it takes
Detection classifiers are cheap, cents per thousand calls, but the engineering effort to design and test an instruction-hierarchy and action-validation layer is substantial and ongoing.