Parallelism here is a property of your executor, not the model: it emits several tool_use blocks in one turn and a serial loop makes them serial anyway.
Why you'd careThe thing you have already noticed
Ask for four things at once and the answer arrives in roughly the time one lookup takes. Ask the same question after a refactor and it takes four times as long, with no change to your prompt. The difference is almost never the model; it is whether your executor awaited each tool call in turn or gathered them. The same stage explains a subtler symptom you have probably shipped without noticing: a comparison table where one row is stale relative to the others, or a summary that quietly covers three of the four things you asked about, because the fourth call failed and its error was one line in a wall of JSON.
In and outWhat goes in, what comes out
| In | One assistant turn whose content array holds N tool_use blocks emitted in a single decode, each with its own id, name and arguments, ending with stop_reason tool_use. The model asserts no ordering between them; they are simply adjacent blocks. |
|---|---|
| Process | The client walks the blocks, schedules them onto a bounded worker pool — asyncio.gather behind a semaphore, a thread pool, a queue — and awaits all of them. Each call carries its own timeout; the turn carries a deadline. Failures are captured as error results, never raised. |
| Out | Exactly one user message whose content array holds N tool_result blocks, one per call id, followed by a single re-prefill of the whole conversation and one new decode. N round trips have collapsed into one. |
The id-to-result binding survives; timing does not. The model sees N results as one simultaneous observation even though they were fetched over a spread of seconds, and nothing records that spread unless you inject timestamps yourself. The fact that the calls ran concurrently is erased too — a parallel run and a serial run produce byte-identical message histories, which is why you cannot tell from a transcript whether your fan-out is actually working.
ConceptThe idea underneath
Decoding is sequential and a tool round trip is expensive on both sides of the wire. Every time the loop returns to the model it must re-prefill the entire conversation so far, decode a fresh assistant turn, then wait on the tool. Serial execution pays that whole stack N times. Fan-out pays it once and then pays only the slowest tool.
Write it as T_serial = N * (t_prefill + t_decode) + sum(t_i) against T_parallel = t_prefill + t_decode + max(t_i). Here t_prefill is the cost of re-reading the growing prompt, mostly served from cache if your cache breakpoints hold; t_decode is the model generating the next set of calls, which happens token by token and does not get cheaper; t_i is tool i's own latency. The second expression is why four lookups feel like one.
The term that bites is max. The maximum of N independent samples is governed by the tail, not the median. If a tool's p99 is 8 seconds, then at N = 4 roughly one turn in twenty-five contains an 8-second call, because 1 - 0.99^4 is about 4 percent. Fan-out converts a rare tail event into a routine one. That is the honest price of the speed-up, and it is why per-call timeouts matter more here than anywhere else in the loop.
Whether a model emits parallel calls at all is learned behaviour, not a setting. Models are post-trained on trajectories containing simultaneous calls and differ substantially in how readily they produce them, and in how well they avoid batching a chain that was never independent.
At a glanceSee it
One decode emits several tool calls; a bounded pool runs them together and one re-prefill resumes.
The knobsHyperparameters and nuance
- disable_parallel_tool_useAnthropic Messages API, set inside tool_choice alongside the type field, default false. Turning it on forces exactly one tool call per turn. It is the right fix for a model that keeps batching dependent steps and the wrong fix for a slow multi-source dashboard.
- parallel_tool_callsOpenAI's equivalent boolean on the request, default true. Same trade under a different name. Setting it false also makes traces trivially diffable, which is worth something while you are evaluating.
- worker-pool widththe semaphore in your executor. Unbounded gather is the common bug: eight tool calls against one host open eight connections at once and trip a per-second limit that serial traffic never touched. A width of 4 to 8 is a sane starting point.
- per-call timeout versus turn deadlinetwo separate numbers, and you need both. With only a turn deadline, one hung call consumes the whole budget and all N results are lost; with only per-call timeouts, N slow-but-legal calls can still overrun the turn.
- connection pool limitshttpx.Limits with max_connections and max_keepalive_connections, defaulting to 100 and 20, or the aiohttp connector equivalent. Sized for serial traffic they silently serialise your fan-out; sized too wide they push the failure downstream into someone else's rate limiter.
EffectHow this stage moves the answer
Fan-out changes what the model is able to know at the moment it decides. Calls issued in the same turn cannot condition on each other, so a task that genuinely needs sequencing — resolve a customer name to an id, then fetch that id's orders — comes back wrong when the model batches it: it guesses the id and reports confidently on a record that does not exist. Partial failure is the other visible change. If one of four calls errors and the error is one line in a wall of JSON, the model will often answer about the three that worked without flagging the gap, so you get a four-column comparison with one silently fabricated column. Consistency degrades too: the results were fetched at different instants, totals do not reconcile, and the model, having no timestamps, presents them as a single snapshot.
EvalsWhat it does to your measurements
The metrics that move are end-to-end latency on multi-fact queries, steps per task, and the downstream 429 rate. The trap is that steps per task is the metric most agent benchmarks report, and fan-out improves it almost for free while tokens per task quietly rises, because a batched agent fetches things it turns out not to need. Optimise steps alone and you will ship a more expensive agent that scores better. Latency benchmarks fail differently: run fan-out against a rate-limited fixture and retries make the parallel path look slower than serial, which is a property of your fixture rather than your design. The Berkeley Function Calling Leaderboard is the standard suite with explicit multiple-call and parallel-call categories, and those are worth reading separately, since a model can be strong at single calls and weak at judging what is genuinely independent.
Failure modesWhen it goes wrong
- A comparison table where one row is stalethe calls completed seconds apart and were presented to the model as one instant, with no per-result timestamp.
- 429 storms the day you enable parallel toolsN calls hit the same host simultaneously against a connection pool and a provider limit both sized for serial traffic.
- The model fetches a record that does not exista dependent chain was emitted as a parallel batch, so call two used an id that call one had not yet returned.
- The API rejects the next request for a missing tool_resultone worker raised instead of returning an error result, so the message went out with fewer results than there were tool_use blocks.
- The context window blows up in a single stepN large results were appended at once, crossing the limit before any per-step budget check had a chance to run.
PapersWhere this comes from
- An LLM Compiler for Parallel Function CallingKim et al., 2023. Plans a task into a DAG of calls with explicit dependencies and executes independent branches concurrently, which is the principled version of what a hand-written gather does, and it reports latency and cost reductions against sequential ReAct-style tool use.
- ReAct: Synergizing Reasoning and Acting in Language ModelsYao et al., 2022 (arXiv 2210.03629). The strictly sequential loop that fan-out is an optimisation of; it matters here as the one-call-per-turn baseline every parallel scheme is measured against.
- Berkeley Function Calling LeaderboardYan et al., 2024. Not a paper in the usual sense but the standard public evaluation, and the one that separates single, multiple and parallel call categories so you can see where a model actually breaks down.