The first call in a fresh process pays DNS, TCP and TLS before the model sees anything, which is why chat UIs open a connection while you are still typing.
Why you'd careThe thing you have already noticed
You have probably seen this: the first request after a deploy takes 900 ms, the next ten take 400 ms, and nothing about the prompt changed. Or you added workers to a batch job and throughput refused to move past a hard ceiling. Neither is the model. Both are this stage — the SDK turning your call into an HTTP request over a connection pool. A cold pool costs a DNS lookup, a TCP handshake and a TLS handshake before a single byte of your prompt is on the wire. A saturated pool makes requests queue inside your own process, invisible to every server-side metric you have. Everything the model does happens after this, and this is where a surprising amount of your p99 is spent.
In and outWhat goes in, what comes out
| In | An application-level call: a list of message objects with roles and content, a model id string, max_tokens, tool definitions as JSON Schema, sampling options, plus credentials and a client-side timeout. Content is still UTF-8 text — nothing has been tokenized. |
|---|---|
| Process | Serialize to JSON, attach headers for authorization, API version and content type, acquire a socket from the pool or pay DNS → TCP → TLS on a cold path, open an HTTP/2 stream, write HEADERS and DATA frames, and arm a deadline plus a retry budget. |
| Out | One request on the wire — typically 2 KB to several hundred KB of JSON carried in HTTP/2 DATA frames — plus a chosen response mode: Accept: text/event-stream for streaming, or a single buffered JSON body for non-streaming. |
Preserved: the exact bytes of your messages and their order, which matters more than it looks because prompt caching keys on those bytes. Lost: your language-level objects. Python and TypeScript values are flattened to JSON, so types, float precision beyond JSON's, and object identity are gone. Also lost is any record of how long the request waited inside your own process, which is why client-side queueing never appears in a provider's latency dashboard.
ConceptThe idea underneath
There is no machine learning in this stage. It is HTTP, and the ideas that matter are the ordinary distributed-systems ones.
The first is connection reuse. A TLS 1.3 handshake on a fresh TCP connection costs a round trip on top of the TCP handshake's own round trip; add DNS resolution and a cold call can spend 200–800 ms before it transmits anything about your prompt. Keeping the socket alive amortizes that across every later request, which is why SDKs ship a connection pool and why warming it before the user finishes typing is a real optimization.
The second is Little's Law, which sets a ceiling almost nobody expects: L = lambda * W. Here L is the number of requests in flight, lambda is the sustainable arrival rate, and W is the average time a request spends in the system. Rearranged, lambda = L / W. If your pool allows 10 concurrent connections and a generation takes 20 seconds, you cannot exceed 0.5 requests per second regardless of how many threads you start. Long generations make W enormous, so a pool size inherited from a REST-API default is almost always wrong.
The third is the streaming decision, which is made here and cannot be revisited later. Setting stream changes the response from one buffered body into a sequence of small events. It does not change the tokens the model produces or the price you pay. It changes when you see them, and it changes how failure presents itself, because the server commits to an HTTP 200 before it knows whether generation will succeed.
At a glanceSee it
One SDK call becoming bytes on the wire, and where a cold pool adds hundreds of milliseconds.
The knobsHyperparameters and nuance
- max_connections / max_keepalive_connectionsthe httpx
Limitsobject the SDK builds its HTTP client with. httpx's own defaults are 100 and 20, but the Anthropic Python SDK overrides them —DEFAULT_CONNECTION_LIMITS = httpx.Limits(max_connections=1000, max_keepalive_connections=100)— so the effective ceiling is far higher than the httpx default you may be picturing. Too low and concurrent calls serialize behind each other, capping throughput at pool size divided by latency; too high and you exhaust file descriptors or trip the provider's per-key concurrency limit. - timeoutthe Anthropic SDKs default to 10 minutes overall (with a 5-second connect timeout), and the units differ by language (Python and Ruby take seconds, TypeScript milliseconds). Too tight and long generations are killed and retried; too loose and one wedged socket holds a worker for ten minutes.
- max_retriesdefault 2 in the Anthropic SDKs, covering 408, 409, 429 and 5xx responses plus connection errors, with backoff. Because timeouts are themselves retried, worst-case wall clock is
timeout × (max_retries + 1)— thirty minutes at defaults. Setting it to zero makes every transient overload user-visible. - streamtrue or false. Non-streaming requests above roughly 16k
max_tokensrisk hitting the client's HTTP timeout, which is why Anthropic's SDKs tell you to stream for large outputs; the TypeScript SDK even scales its default timeout upward for large non-streaming requests. This is provider- and SDK-specific behaviour. - HTTP versionHTTP/2 multiplexes many requests over one connection, so pool size and concurrency decouple. On HTTP/1.1 each in-flight request needs its own socket, and the pool becomes the concurrency limit directly.
- TCP_NODELAYdisables Nagle's algorithm. Left enabled, a small write can sit up to about 40 ms waiting to coalesce with the next one, which is pure added latency on a request that fits in a single segment.
EffectHow this stage moves the answer
Transport does not change a single token, and then it changes the answer anyway, in three observable ways. A read timeout shorter than your longest generation does not fail randomly, it fails selectively, killing exactly the long, thorough answers and letting short ones through. Your production output distribution ends up shorter than the model's, and nobody notices because the losses are logged as infrastructure errors rather than quality ones. Second, an automatic retry after a timeout re-runs a request the server may already have completed; with any sampling randomness the second answer differs from the first, and for a tool-using agent it can re-execute side effects. Third, an intermediary that buffers a response can truncate it at a size limit, so your JSON parser fails on output the model finished correctly.
EvalsWhat it does to your measurements
Latency benchmarks are where this stage does the most damage. A harness that constructs a fresh client per sample pays a cold handshake on every request, so measured time-to-first-token is inflated by hundreds of milliseconds production never sees, and the inflation is bimodal, which wrecks percentiles. Run at concurrency and the opposite happens: a pool smaller than your worker count turns client-side queueing into what looks like server latency, and p99 becomes a measurement of your own thread pool. Retries are worse than either because they are silent. A harness with two automatic retries reports a zero percent error rate against a fleet that was shedding a third of its load, while token usage and your bill reflect attempts you never counted. Warm the client, pin concurrency explicitly, and log retry counts alongside latency.
Failure modesWhen it goes wrong
- The first request after process start is 300–800 ms slower than every request after itDNS lookup plus TCP and TLS handshakes on a cold connection pool.
- Throughput plateaus at exactly N concurrent requests however many workers you add
max_connectionson the HTTP client is the real limit, and Little's Law caps you at pool size divided by request latency. - Timeout errors that correlate with prompt complexity rather than server loada fixed read timeout set below the generation time of your longest answers.
- Duplicate tool side effects or double-billed tokensthe SDK retried a request the server had already completed, and nothing in the transport tells the client which.
- Streaming works locally and arrives in one lump in productiona reverse proxy or CDN is buffering the response body; nginx with
proxy_buffering onis the usual culprit.
PapersWhere this comes from
There is no research literature for this stage. It is HTTP, and the authoritative documents are specifications and reference implementations rather than papers: RFC 9113 defines the HTTP/2 framing and stream multiplexing your SDK relies on, RFC 8446 defines the TLS 1.3 handshake and its round-trip cost, and RFC 896 (John Nagle, 1984) is still the reason small writes stall. One systems paper is directly relevant:
- The Tail at ScaleJeffrey Dean and Luiz Andre Barroso, 2013. Established that tail latency in request-serving systems is dominated by rare slow responses rather than the median, and proposed hedged and tied requests as the mitigation; it is why a client-side deadline and retry policy belong at this stage rather than deeper in the stack.