The model never runs anything; it emits a JSON request and your dispatcher decides whether to honour it, which is where every real security property actually lives.
Why you'd careThe thing you have already noticed
You have watched a chat sit motionless for thirty seconds, spinner turning, and then answer with "I checked the deployment logs and...". Nothing was checked by the model during those thirty seconds; it had stopped generating before the wait even began. What was running was your own process: resolving a tool name, validating arguments against a JSON Schema, opening a subprocess or an HTTP connection, and waiting on a timeout you configured. The same stage explains the opposite symptom, where a tool returns in 8 ms with a permission error and the model then narrates a result it never received. Tool dispatch is ordinary software on your side of the boundary, and it fails in ordinary software ways.
In and outWhat goes in, what comes out
| In | A decoded tool_use block: an opaque call id, a tool name, and an arguments object the model produced token by token. Out of band, the dispatcher also holds the caller's identity, session scope, granted permissions and any per-tenant budget. |
|---|---|
| Process | Registry lookup on the name, JSON Schema validation and type coercion of the arguments, an authorization check against the caller rather than against the model, then execution inside a bounded context — a subprocess with seccomp, a container namespace, a gVisor or Firecracker guest, or a remote MCP server — under a wall-clock timeout, a memory cap and an egress allowlist. |
| Out | A tool_result block carrying stdout, a JSON document, base64 bytes or an image, truncated to a byte or token cap; or an is_error result with a message string. Separately, telemetry the model never sees: duration, exit code, bytes read, and every policy decision taken. |
What survives is the call id, the only thing binding a result back to the request that produced it, plus whatever text fits under the cap. What is lost is everything the sandbox observed but did not return: exit codes flattened into a generic error string, stderr discarded, the middle of a large output silently removed. The model reasons over the surviving text alone, so a truncated result and a genuinely short one are indistinguishable to it.
ConceptThe idea underneath
There is no code path inside the network that calls anything. A tool is a string in the prompt — a name, a description, a JSON Schema — and the model's entire participation is emitting tokens that spell that name and a plausible arguments object, frequently under a grammar mask that forces the JSON to parse. Function calling is a learned output format, not an execution capability. The behaviour comes from post-training on trajectories where producing that format was rewarded.
So this stage has almost no machine learning in it, and pretending otherwise is how people build insecure agents. The idea underneath is the reference monitor from operating-system security: a component that mediates every request for a resource, cannot be bypassed, and is small enough to audit. Two of its classical properties matter here. Complete mediation means the check runs on every call, not once at session start. Least privilege means the sandbox gets exactly the capabilities the tool needs, because the model's effective privilege is not the model's at all — it is whatever credentials your dispatcher process happens to be holding when it fires.
Isolation is a ladder, and each rung trades startup latency for a smaller trusted computing base. Running the tool in-process gives you none: one path traversal reads your service account key. A subprocess behind a seccomp filter blocks most syscalls and costs roughly a millisecond to start. A container adds namespaces and cgroups. A user-space kernel such as gVisor intercepts syscalls before they reach the host. A microVM such as Firecracker gives the tool its own guest kernel for an advertised boot under about 125 ms. Choose the rung by blast radius, not by what is convenient to wire up.
At a glanceSee it
Where a tool call crosses from model tokens into your process, and what confines it there.
The knobsHyperparameters and nuance
- per-tool timeoutthe wall-clock ceiling on one call. The MCP TypeScript SDK's request timeout defaults to 60 seconds; Claude Code exposes MCP_TOOL_TIMEOUT for the same purpose. Too short and legitimately slow tools (a full build, a cold warehouse query) fail every time; too long and one hung call holds the entire agent turn until the outer deadline.
- result size capthe byte or token limit on what a tool may return, for example MAX_MCP_OUTPUT_TOKENS in Claude Code. Too small and the model draws conclusions from a head-only slice; too large and a single grep consumes a third of the context window and triggers compaction two steps later.
- --network=none and the egress allowlistDocker's flag or an equivalent proxy rule. Default deny is the only defensible posture for tools that execute model-authored code. Allow everything and exfiltration is a one-line tool call; allow nothing and any tool that legitimately fetches breaks, so allowlist by host rather than toggling the internet on and off.
- --memory and --pids-limitthe cgroup caps. Without them a fork bomb or a runaway dataframe read takes down the host that also serves your API. Set too tight and ordinary workloads get OOM-killed at a size the model has no way to predict, producing an error it cannot act on.
- sandbox lifetimeper-call versus per-session. Per-call is safer and pays container start on every tool use; per-session keeps working state such as an installed package or a loaded dataset, and lets one poisoned call contaminate every later one.
- concurrency limitthe semaphore width on simultaneous sandboxes. Too wide and N agents times M tools exhausts host memory; too narrow and tool calls queue invisibly, surfacing as latency you will misattribute to the model.
EffectHow this stage moves the answer
Every property of this stage reaches the user as a change in what the answer claims. A timeout that fires returns an error string, and the model's most likely continuation after an error is either an apology with no content or a confident summary assembled from its priors — which is how you get plausible numbers that were never fetched. A result cap that truncates without a marker is worse: the model reads a complete-looking document and states conclusions about the part it never saw, with no hedging, because nothing in its input signalled incompleteness. A sandbox with no egress turns "the API is unreachable from here" into "there are no results", a claim about the world rather than about the network. And raw stack traces in a tool_result reliably derail the turn: the model starts debugging your dispatcher instead of answering the question.
EvalsWhat it does to your measurements
Tool error rate by class and p95 tool latency both live here, but the way this stage corrupts a measurement is subtler. Sandbox reuse warms everything: the second run of a benchmark hits a container with the dependency already installed and the page cache already hot, so your p50 looks excellent and the cold-start tail that breaks production never appears in the numbers. Agent benchmarks are worse, because harness failures score as model failures. On tool-use suites such as tau-bench, a task marked failed is frequently a 60-second timeout, a truncated tool result or a rate-limited fixture rather than a reasoning error. Report tool error rate alongside task success, split by whether the error came from the tool, the sandbox or the model's arguments. Two harnesses with different result caps will rank the same models differently.
Failure modesWhen it goes wrong
- The model narrates a result it never receivedthe dispatcher swallowed an exception and returned an empty string instead of an is_error block, so an empty success is what reached the context.
- Works in dev, times out in prodproduction creates a sandbox per call and container start is billed against the tool's own timeout, so the clock expires before the work begins.
- One slow tool holds the entire turnonly a global deadline exists; with no per-tool timeout the slowest call defines the user-visible latency of the whole step.
- Answers cite only the first part of a large filethe result was truncated to fit a token cap with no truncation marker, so the model treated a prefix as the whole.
- An injected document causes a real side effectthe dispatcher validated the tool name and schema but authorised against the model's arguments rather than the caller's permissions, so mediation was incomplete.
PapersWhere this comes from
- ReAct: Synergizing Reasoning and Acting in Language ModelsYao et al., 2022 (arXiv 2210.03629). Established the interleaved thought, action and observation loop that every tool dispatcher still implements; it is why a tool result re-enters the model as an observation in the transcript rather than as a function return.
- Identifying the Risks of LM Agents with an LM-Emulated SandboxRuan et al., 2023. Built ToolEmu, which emulates tool execution so risky agent behaviour can be measured without real side effects, and found that even strong models take damaging actions at non-trivial rates. It is the empirical case for confining execution rather than trusting the caller.
- Firecracker: Lightweight Virtualization for Serverless ApplicationsAgache et al., NSDI 2020. Not an ML paper: it showed that a purpose-built microVM can deliver per-tenant kernel isolation at close to container startup cost, which is what makes per-call VM sandboxing practical for tool execution at all.
- Model Context Protocol specificationAnthropic, 2024. Also not research; it is the wire contract most dispatchers now speak, and the place where client-side responsibilities such as timeouts, user consent and tool result shape are actually written down.