Home › Security › Zero trust for agents
🔒 · Operate

Zero trust for agents

Treating everything an agent reads, including tool output, as untrusted input.

In one line

An agent's own tool results are attacker-controlled text, so the boundary is not around the agent but around every value that enters its context.

ConceptWhat it is

Zero trust is the principle that no network position confers trust — every request is authenticated and authorised on its own. Applied to agents it becomes something sharper: no text confers trust either, whatever produced it.

The reason is that an agent cannot reliably distinguish instruction from data. A web page it fetched, a document it retrieved, a tool's error message — all of it arrives as tokens in the same context as the system prompt, and any of it can attempt to redirect the agent. The agent is not the perimeter; it is the thing being defended.

How it worksThe mechanics

Each value entering the context is tagged by origin and is never treated as carrying authority. Instructions come only from the system prompt; everything else is content to be reasoned about. Where a tool result could plausibly contain instructions, it is delimited and labelled as untrusted rather than concatenated in.

The second half is that authority lives outside the model. A tool call that matters is checked against a scope the model cannot widen, and irreversible actions are gated on a confirmation the model cannot forge. The assumption is that the agent will eventually be talked into attempting the wrong thing, and the design makes that attempt fail.

At a glanceSee it

Zero trust for agents diagram

Authority comes only from the system prompt, and enforcement lives outside the model. The design assumes the agent will be talked into trying something it should not.

When to use itWhere it fits

  • Any agent that reads content it did not author — the web, user uploads, retrieved documents, other agents.
  • Any agent with a tool that writes, sends, pays or deletes.
  • Multi-agent systems, where one agent's output is another's untrusted input and the chain hides the origin.
  • Wherever a successful injection would be expensive, which is the honest test rather than whether one is likely.

When NOT to use itLimits & anti-patterns

  • A closed loop over content you fully control, where the ceremony buys little.
  • As a claim of immunity — this bounds the damage of injection, it does not prevent injection.
  • As a reason to skip input filtering; the two are complementary and filtering catches the cheap attempts.
  • When it degenerates into confirming every action, which trains the human to click through and removes the control.

Trade-offsAdvantages & costs

Advantages
  • Bounds blast radius by construction rather than by hoping the prompt holds.
  • Survives model upgrades, because it does not depend on any particular model resisting persuasion.
  • Makes the dangerous set explicit — the tools requiring confirmation are a short, reviewable list.
  • Composes with IAM scoping, so two independent controls have to fail together.
Trade-offs & costs
  • Confirmation steps cost the autonomy that motivated building an agent in the first place.
  • Tagging provenance through a long context is fiddly and easy to lose across a summarisation step.
  • Delimiting untrusted content helps but is not a hard boundary; models still sometimes follow it.
  • The discipline is invisible when it works, so it is the first thing cut under delivery pressure.

ExampleIn the real world

A research agent is asked to summarise competitor pricing and fetches a page containing, in white text, an instruction to email its context to an outside address. The agent has an email tool. Under zero trust the fetched page is labelled untrusted content, the email tool is out of scope for a research task, and the attempt fails at the scope check and appears in the log. Nothing about the model changed; the boundary was never inside it.

ToolsHow to implement it

  • Scoped, per-task tokensthe enforcement point, since it is the one thing the model cannot argue with.
  • Structured tool schemas with strict validationarguments checked against types and allowlists before execution.
  • Human-in-the-loop confirmation for irreversible callsreserved for a short list, so it stays meaningful.
  • Content delimiting and provenance tagscheap, imperfect, and worth doing as one layer among several.

Cost & effortWhat it takes

Delimiting and tagging cost a few hundred tokens per turn. Confirmation costs human time, which is the real budget line and the reason the gated list must stay short. Engineering effort is moderate and front-loaded; retrofitting it onto a shipped agent with broad tool access is the expensive path.

What changedWhat changed here

Written inYou approved this and it changed the page
  • Updated this page A research framework for runtime governance of agent tool actions, with fail-closed execution, that supports the page's zero-trust stance.

    arXiv cs.AI · 18 Aug 2026 · source

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning