Home › What Happens After You Hit Enter › Prefix-Aware & Session-Sticky Routing
Pipeline stage · Operate

Prefix-Aware & Session-Sticky Routing

The load balancer picks the replica that already holds your prompt's prefix — or your conversation's history — in GPU memory, instead of the least busy one.

In one line

A cache hit is a property of which machine you land on, so the same prompt can cost a full prefill on one replica and almost nothing on another.

Why you'd careThe thing you have already noticed

Your system prompt has not changed in weeks, but your cached-token percentage moves by tens of points day to day and first-token latency for the same request varies by seconds. Turn seven of a chat is sometimes slower than turn six. The usual suspects — prompt content, model version, overall load — are all constant. What varies is placement. KV cache lives in one replica's GPU memory, not in a shared tier, so reuse is only possible when the load balancer sends you back to the machine that already holds your prefix. A round-robin balancer throws that away on every request. This stage is the balancer being told to care.

In and outWhat goes in, what comes out

InBlock-granular hashes of the prompt prefix, or a session id, plus a periodically refreshed summary of what each replica currently holds — a radix tree of token blocks or a compact digest — and each replica's live load in queue depth and running sequences.
ProcessFor each replica, compute the longest prefix match in blocks and score it against load, roughly score = w_affinity * matched_tokens - w_load * queued_tokens, then take the maximum. A session already bound to a replica may be rebound if that replica is saturated.
OutA target replica, an expected count of reusable KV blocks, and a binding decision. A rebind means the KV built on the previous replica is stranded and will be evicted without ever being read again.

The request content is untouched. What is lost is placement independence: after this stage, latency and cost depend on routing history rather than on the request alone, and two identical calls are no longer interchangeable. A rebind discards real work — the prefill that produced the old replica's blocks is gone, and the new replica recomputes from token zero, because a KV prefix cannot be transplanted without physically moving the bytes.

ConceptThe idea underneath

The mechanism to understand is why cache reuse is a prefix property and never a substring one.

Attention at position i reads the key and value vectors of every position before it, so the K and V stored for token i depend on every token from 0 to i. Change one byte at position 3 and every cached vector after it is wrong. That is why caching is implemented as a chain of block hashes: h_i = hash(h_{i-1}, tokens_of_block_i), with a fixed block size (16 tokens is the vLLM default). Two requests share KV exactly up to the first block where their hashes diverge. A timestamp at the top of your system prompt does not reduce your hit rate, it zeroes it.

SGLang's RadixAttention organizes those blocks as a radix tree per replica, so many requests sharing a system prompt share one physical copy of its KV and the reusable length for a new request is a tree lookup. Skipping that prefill is where the saving lives: prefill is compute-bound and grows with prompt length, while decode is memory-bandwidth-bound and barely depends on it, so a long shared prefix is nearly all of the recoverable work.

Routing on top of that is a classic cache-affinity load-balancing problem — consistent hashing with bounded loads, except the cache is tens of gigabytes of GPU memory that turns over every few hundred milliseconds. Pure affinity concentrates the most popular prefix on one replica and melts it. The standard fix is a bounded rule: prefer the cache owner unless its load exceeds the fleet mean by more than some factor, then fall back to least-loaded.

At a glanceSee it

Prefix-Aware & Session-Sticky Routing diagram

Prefix hashes compared against each replica's cache, with load overriding affinity when the owner runs hot.

The knobsHyperparameters and nuance

  • enable_prefix_caching(vLLM) — the master switch. On recent versions it defaults on, but a config that turns it off makes every other knob here inert and every routing decision pointless.
  • block_size(vLLM, default 16) — the hashing and allocation granularity. Larger blocks mean fewer lookups and coarser matching, so a prefix that diverges mid-block loses the whole block; smaller blocks mean more bookkeeping per token.
  • affinity versus load weightthe balance term in cache-aware routers, exposed in some form by the vLLM production-stack router and SGLang's front end. All affinity gives spectacular hit rates and one melted replica; all load gives even utilization and no cache.
  • minimum match lengthhow many matched tokens are required before affinity may override load. Set it near zero and you route on the coincidence that two prompts begin with the same chat template.
  • cache-state sync intervalhow often replicas report what they hold. Long intervals send requests to replicas that already evicted the blocks; short intervals cost chatter and still lag reality.
  • session TTLhow long a conversation stays pinned to a replica. Too long and dead sessions pin capacity; too short and a multi-turn chat loses its history's KV between turns.

EffectHow this stage moves the answer

Most of what this stage changes is cost and latency, and it would be dishonest to claim otherwise. But there are three routes to the answer itself. First, numerics: reusing a cached block computed under a different batch shape can produce values that differ in the last bits from what a fresh prefill would compute, and at temperature zero a near-tie between two candidate tokens can flip on that difference — so the same prompt can produce different text on a cache hit than on a miss. Second, adaptive products degrade under a miss: if a slow first token triggers a fallback to a smaller model or a reduced reasoning budget, routing has chosen your answer indirectly. Third, in stateful serving where the session id also keys server-side conversation or tool state, a rebind under load can drop that state, and the assistant appears to forget the conversation mid-thread.

EvalsWhat it does to your measurements

Two numbers move — time-to-first-token and cached-token ratio — and both are trivially faked by a careless harness. A benchmark that loops over the same 200 prompts warms every replica and measures a steady state production never reaches; report cold-start TTFT separately. A harness with a single client against a single endpoint lands everything on one replica and never exercises the routing policy at all. Measure hit rate as cached input tokens divided by total input tokens, not as a fraction of requests: one long cached system prompt and one short uncached one are not equal. Note also that a warm cache can change numerics, so a rerun on a warmed fleet is not bit-identical to the first run even at temperature zero — do not chase that as a scoring bug. Finally, a fresh cache entry is not readable until the request that wrote it begins streaming, so N identical requests fired in parallel all miss.

Failure modesWhen it goes wrong

  • Cache hit rate collapses to near zero right after a deployreplicas restarted with empty caches, and any edit to the system prompt invalidates every stored prefix as well.
  • Hit rate is near zero even though the system prompt is identical across requestssomething volatile sits at the top of the prompt: a timestamp, a request id, a user name, or a tool list serialized in non-deterministic order.
  • One replica pinned at full utilization while the rest idleunbounded affinity routing with a single dominant shared prefix.
  • Latency gets worse immediately after scaling outthe new replicas have cold caches, and a load-weighted router preferentially sends them traffic precisely because they are idle.
  • A conversation loses its context after a period of inactivitythe session was rebound to a replica that never held its history, in a stack that treated the session id as if it carried state.

PapersWhere this comes from

  • SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng et al., 2024 (arXiv:2312.07104). Introduced RadixAttention, which stores KV blocks in a radix tree so requests sharing a prefix share one copy and reusable length becomes a cheap lookup; this is the data structure a prefix-aware router queries.
  • Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon et al., 2023 (arXiv:2309.06180). Established block-granular, non-contiguous KV allocation, which is the precondition for sharing a prefix across requests at all.
  • Prompt Cache: Modular Attention Reuse for Low-Latency InferenceIn Gim et al., 2024. Explores reusing attention state for segments that are not at the start of the prompt, and shows it requires re-encoding positions and remains an approximation; useful mainly as evidence for why ordinary caching is strictly prefix-only.
  • Consistent Hashing with Bounded LoadsVahab Mirrokni, Mikkel Thorup and Morteza Zadimoghaddam, 2016. A general result on capping how far above average any cache-affine node may be loaded before traffic spills elsewhere; the standard formalization of the affinity-versus-load tension this stage manages.
A living map of modern AI — kept current every morning