Refusing a request in five milliseconds is kinder than accepting it and making every request already running twenty seconds slower.
Why you'd careThe thing you have already noticed
You have seen the error that says overloaded, please retry, and read it as the system failing. It is the system working. The alternative — accepting everything — is a failure mode you have also seen: a two-minute traffic spike followed by latency that stays broken for twenty minutes after the spike ends, because every request that timed out was retried and the queue never drained. That is congestion collapse, and admission control is what prevents it. This stage reads queue depth, free KV-cache blocks and projected time-to-first-token, and decides whether accepting your request would break the promise the system already made to the requests it is running.
In and outWhat goes in, what comes out
| In | Live engine state — waiting-queue depth, free KV-cache blocks, number of running sequences, prefill backlog — together with this request's prompt length, its max_tokens, and its tenant or priority class. |
|---|---|
| Process | Estimate the memory this request will occupy at completion, roughly prompt_tokens + max_tokens multiplied by bytes-per-token, and compare against a free-block watermark. Estimate queueing delay from the current batch and compare against the time-to-first-token target. Then admit, hold or shed. |
| Out | One of three outcomes: admitted into the engine's waiting queue with a scheduling priority; held in a bounded queue; or rejected with HTTP 429 (rate_limit_error) or 529 (overloaded_error) plus a retry-after header telling the client when to return. |
Nothing about your request is altered — a shed request and an admitted one are byte-identical. What is irreversibly lost is optionality. Once admitted, a request holds KV blocks, and the only way to reclaim them is preemption, which throws away completed prefill work and pays for it twice. Shedding costs microseconds; preempting costs whatever prefill already ran. That asymmetry is why this decision sits in front of the engine rather than inside it.
ConceptThe idea underneath
This is systems engineering, but systems engineering driven by one number that comes straight out of the model architecture: how much memory a single token of context costs.
Every token in the KV cache stores a key vector and a value vector at every layer, for every key/value head. So bytes_per_token = 2 * n_layers * n_kv_heads * d_head * bytes_per_element, where the leading 2 is K and V. For a 70B-class model with 80 layers, 8 grouped-query KV heads, head dimension 128, in 16-bit precision, that is 2 × 80 × 8 × 128 × 2 = 327,680 bytes, about 320 KiB per token. A 100,000-token context is roughly 32 GB of GPU memory for one request. That arithmetic, not compute, is why concurrency limits feel so low.
Compute is elastic in a way memory is not. You can make a batch slower; you cannot make a KV block smaller. Admission control is therefore a memory-admission problem wearing a latency costume.
The second idea is goodput versus throughput. Throughput counts requests finished per second. Goodput counts requests finished per second within their latency target. A queue that never sheds maximizes throughput and can drive goodput to zero, because everything completes eventually and nothing completes in time.
The third is why unbounded queues are dangerous. When clients retry on timeout, offered load rises as service quality falls, which raises load further. The system settles into a stable bad state that outlives the trigger that caused it — a metastable failure. Bounded queues, shedding with a retry-after, and a client-side retry budget are the standard escape.
At a glanceSee it
How the scheduler decides whether accepting one more request would break the ones already running.
The knobsHyperparameters and nuance
- gpu_memory_utilization(vLLM) — fraction of VRAM given to weights plus KV cache, default 0.90. Too high and a peak-length request triggers an out-of-memory error that kills the whole engine; too low and the KV pool is small, so the scheduler preempts constantly under ordinary load.
- max_num_seqs(vLLM) — hard cap on sequences running concurrently. The V1 engine defaults to 1024; the 256 widely quoted in older material was the V0 default and no longer applies. Set above what KV can hold at realistic lengths and you get preemption thrash; set far below and the GPU idles between decode steps.
- max_num_batched_tokens(vLLM) — token budget per scheduler step, which with chunked prefill decides how much prefill work can crowd out decode. Too small and long prompts take many steps to admit; too large and one big prefill stalls every streaming response in the batch.
- max_model_lenthe longest context the engine accepts. A longer request is rejected at admission rather than silently truncated, which is the behaviour you want but not always what a proxy in front of the engine does.
- retry-after and the client retry budgetshedding only helps if clients honour it. The Anthropic SDKs default to two retries with backoff and read the
retry-afterheader; a hand-rolled client that retries immediately converts shedding into a retry storm.
EffectHow this stage moves the answer
The visible effect is not a worse answer, it is a different answer path. Applications almost always have a fallback for a shed request — a smaller model, a cached response, a canned apology — so the moment your traffic peaks is exactly the moment users get your second-best output. Preemption is stranger. When the engine reclaims KV blocks from a running sequence and re-prefills it later, the regenerated continuation is sampled again; with any randomness in sampling, the text after the preemption point can diverge from what was already streamed, which reads as the assistant contradicting itself mid-paragraph. And any policy that estimates cost from prompt length plus max_tokens systematically disfavours long requests, so under sustained load your product quietly stops producing long-form output at all.
EvalsWhat it does to your measurements
This stage sets your goodput ceiling, and goodput is the metric almost no benchmark reports. A harness running at fixed concurrency 8 will never trigger shedding and tells you nothing about the same system at concurrency 800; the interesting behaviour only exists above the watermark. Worse, most harnesses retry on 429 by default, which converts a capacity failure into latency — you get a zero percent error rate and a p99 that is mostly your own backoff sleep. If the harness also has a model fallback, an accuracy regression can appear that has nothing to do with the model you thought you were measuring. Report load tests as a curve of completed-within-target against offered load, and log shed counts, retry counts and preemption counts next to latency.
Failure modesWhen it goes wrong
- Overload errors while GPU utilization sits around 60 percentthe watermark is computed from prompt length rather than actual free blocks, so the admitter reserves memory for a worst case that never arrives.
- Latency spikes during a burst and never recovers after the burst endsan unbounded queue plus client retries; the system has entered a metastable state where retry load sustains the overload that caused the retries.
- Excellent throughput numbers and angry usersrequests are completing, just not within anyone's deadline; throughput was optimized where goodput was the objective.
- Streaming responses that restart or contradict themselves mid-answerthe scheduler preempted a running sequence and recomputed it, re-sampling the continuation.
- All the rejections land on one tenantthe shed order is arrival-based rather than fairness-based, so the burstiest tenant absorbs the entire rejection budget.
PapersWhere this comes from
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM ServingRuoyu Qin et al., 2024. Describes a production serving stack that predicts whether an incoming request can meet its latency target before admitting it and rejects early rather than queueing; the clearest published account of prediction-based load shedding for LLM inference.
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong et al., 2024. Makes goodput — requests meeting per-phase latency targets — the optimization objective instead of raw throughput, which is precisely the distinction admission control exists to enforce.
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal et al., 2024. Introduced chunked prefill and stall-free scheduling, showing that how you admit a long prompt determines whether every already-running decode stalls behind it.
- Metastable Failures in Distributed SystemsNathan Bronson et al., HotOS 2021. Named the failure mode in which retries sustain an overload after its original trigger has passed; it is the general argument for shedding rather than queueing.