Prompt injection is malicious text that tricks a model into obeying an attacker instead of you.
DefinitionWhat it means
Prompt injection is an attack in which text embedded in a user input, a retrieved document, or a webpage the model reads contains instructions designed to override the system's intended behavior, causing the model to leak data, ignore rules, or take unintended actions.
Why it mattersWhy you should care
Any agent that reads untrusted content, emails, web pages, tool outputs, is exposed to prompt injection, so shipping such a product requires layered defenses like input sanitization, permission scoping, and output monitoring rather than trusting the model alone to resist manipulation.
At a glanceSee it
Two roads into one hijack — direct attacks the user types, and indirect ones smuggled through a poisoned page, email, or tool output — both end with the model obeying the attacker.
A layered defense quarantines untrusted content as data, sandboxes tools to least privilege, and gates risky actions behind a human — skipping that boundary is precisely what lets an injection execute.
Where you see itIn the wild
- Security disclosures against browser and email agents that read attacker-controlled pages.
- Guardrail products that scan tool outputs before feeding them back to the model.
- Security reviews on defending an agent against injected instructions.
What changedWhat changed here
- Breaking Claude Code Opus 5 Auto Mode
A prompt-injection researcher found an attack on Claude Code's default auto mode that he claims works 80% of the time, using a zip archive to hijack a base64 import. For anyone running agents in auto-approve mode, treat it as untrusted-input territory and keep human approval in the loop for downloads and cleanup commands.
- Auto mode is now the default in Claude Code for Pro, Max, and Team plans
Anthropic is making auto mode the default for new Claude Code sessions on Pro, Max, and Team plans starting August 14, letting the agent plan and execute multi-step coding tasks with less human sign-off. If you build on Claude Code, expect a more autonomous agent workflow — and re-check your prompt-injection guards before the switch.
Three kinds of claim, strongest first. Signal runs every morning.