Home › What Happens After You Hit Enter › Controlling speculation under load
Pipeline stage · Operate

Controlling speculation under load

Knowing when to stop guessing — speculative decoding helps an idle GPU and hurts a busy one.

In one line

Speculative decoding spends spare arithmetic to save memory reads, so it only wins while the GPU is memory-bound — a deep batch has already spent that headroom.

Why you'd careThe thing you have already noticed

You turn on speculative decoding, measure it the way everyone measures it — one request, one user — and watch decode go nearly twice as fast. You ship it. Under real traffic, p50 latency is flat and total throughput has dropped maybe fifteen percent. Nothing is broken and nothing in the logs looks wrong. What changed is batch size. Speculation buys latency by doing extra arithmetic on tokens that may be discarded, and that trade is only free while the GPU is sitting idle waiting on memory. Once the batch is deep enough to keep the tensor cores busy, draft tokens stop being free and start competing with other users' real tokens for the same FLOPs.

In and outWhat goes in, what comes out

InPer-step serving state: current running batch size, the rolling acceptance rate over recent steps, queue depth and waiting requests, plus the configured drafting method — a small draft model, extra decoding heads, or an n-gram lookup over the prompt.
ProcessDraft k tokens cheaply, run one target-model forward pass over the k+1 candidate positions in a single batched call, compare the target and draft distributions, accept the longest verified prefix under the rejection rule, and discard the rest. A controller decides k, or decides to skip speculation entirely this step.
OutBetween 1 and k+1 accepted tokens for one target forward pass, an updated acceptance-rate estimate feeding the controller, and KV blocks for the rejected drafts released back to the allocator.

The output distribution is preserved. The rejection scheme of Leviathan et al. and Chen et al. is provably distribution-matching — with exact arithmetic, speculative decoding samples from the same distribution as the target model alone, which is what makes it a pure optimisation rather than a quality trade. What is lost is byte-level reproducibility against a non-speculative run, because the arithmetic path differs, and any assumption of steady per-token latency: output now arrives in bursts of accepted tokens, so inter-token gaps become bimodal.

ConceptThe idea underneath

There is no neural network idea here. This is a roofline argument, and it is worth doing the arithmetic once because every intuition about speculation follows from it.

Decode is memory-bound. To produce one token you read every weight in the model from HBM and do roughly two FLOPs per parameter. With bf16 weights that is two bytes read per parameter, so arithmetic intensity is about 1 FLOP per byte at batch size 1. An H100 has roughly 990 bf16 TFLOPS against roughly 3.35 TB/s of HBM bandwidth — a machine balance near 300 FLOPs per byte. At batch 1 you are using well under one percent of the arithmetic you paid for. Batching raises intensity linearly: batch B gives about B FLOPs per weight byte, so you would need a few hundred concurrent sequences to saturate the tensor cores, and KV cache capacity usually stops you well short.

Speculation exploits the same slack from a different direction. Instead of amortising the weight read over B different users' tokens, amortise it over k+1 candidate positions from one user. Expected accepted tokens per target pass, with per-token acceptance probability alpha and draft length k, is (1 - alpha^(k+1)) / (1 - alpha) — a geometric sum because acceptance stops at the first rejection. At alpha = 0.7, k = 4 that is about 2.9 tokens per pass; at alpha = 0.4 it is about 1.6.

Acceptance rate is a property of the workload, not of the model. Code and structured output draft extremely well. Free-form creative prose drafts badly. N-gram drafting, which copies from the prompt, is near-free and excellent for summarisation and edit tasks where the output quotes the input, and close to useless anywhere else.

At a glanceSee it

Controlling speculation under load diagram

Draft tokens are verified in one target pass while a load controller disables the scheme as batches grow.

The knobsHyperparameters and nuance

  • num_speculative_tokens(vLLM, inside --speculative-config) — the draft length k, typically 3–5. Too small and you barely amortise the weight read; too large and verify cost grows while the geometric acceptance sum flattens, so rejected work dominates once acceptance drops below roughly 0.6.
  • method(ngram, ngram_gpu, suffix, eagle / eagle3, mtp, medusa, mlp_speculator, draft_model) — n-gram and suffix drafting cost no GPU memory but only pay off when the output copies the input or a cached earlier response. A draft model consumes VRAM that would otherwise hold KV blocks, shrinking your maximum batch and changing the regime you tuned for.
  • num_speculative_tokens_per_batch_size(vLLM) — the batch-aware control, and the setting that decides whether speculation helps or hurts in production. It takes a schedule of (range_start, range_end, num_speculative_tokens) triples over inclusive batch-size ranges, so you shorten or zero out the draft as concurrency rises instead of paying verify cost you cannot amortise. The knee is typically somewhere between 4 and 32 running sequences depending on hardware and model size — measure it. Note that this replaced an older on/off switch: speculative_disable_by_batch_size, later disable_by_batch_size, belonged to the V0 engine and disappeared when V0 was removed in vLLM v0.25.0, so copying it out of an older runbook now fails at startup.
  • --max-num-seqs and --max-num-batched-tokens(vLLM) — these determine where the batch-size knee actually sits. Tuning speculation without reference to them means tuning against a batch distribution you do not have.
  • rejection_sample_method / draft_sample_method(vLLM) — verification in current vLLM is a strict, distribution-preserving rejection sampler. The relaxed typical acceptance criterion from the Medusa line of work is real in the literature but is no longer available in vLLM: it and the --spec-decoding-acceptance-method flag went with the V0 engine. Treat that trade as something you would have to implement, not something you can switch on. The accepted values for rejection_sample_method have drifted between the shipped docs table and the config source, and one of them is a synthetic mode meant for benchmarking a target acceptance rate rather than for serving — read the version you are actually running before setting it.
  • gpu_memory_utilization(vLLM) — indirect but decisive: draft weights come out of the same budget as the KV cache, so raising speculation lowers concurrency, which raises the odds speculation was worth it, which is a feedback loop you have to measure rather than reason about.

EffectHow this stage moves the answer

Under the strict rejection rule, speculation should not change the answer at all, and mostly it does not. What it changes is the shape of delivery: users see three tokens land at once, then a pause, then four more. In a UI that animates per token, that reads as stuttering rather than as speed, and the fix belongs in the client's render loop, not the server. Where it does change output, there are two causes. Floating-point differences between a k+1-position verify pass and a plain single-position decode move logits in the last bits, so near-ties flip and a byte-exact regression test against a non-speculative baseline will fail even though nothing is wrong. And if your engine defaults to a relaxed or typical acceptance mode rather than the exact rejection rule, output quality genuinely shifts — more accepted drafts, slightly off-distribution text. Check which mode is active before you attribute a quality change to a model update.

EvalsWhat it does to your measurements

Never benchmark speculation at batch size 1, which is where every quick test lands by default. The numbers that matter are acceptance rate, mean accepted length, and goodput — total tokens per second across all concurrent users, not tokens per second down one stream. Four specific ways a measurement lies. A warm prefix cache inflates the apparent speedup because prefill disappears from the timing and decode is all that is left. Acceptance rate measured on a code benchmark does not transfer to a chat workload, so a single tuned k is wrong for half your traffic. Time-to-first-token is unaffected by speculation, so a harness that reports only TTFT will show neither the benefit nor the harm. And a load generator running at fixed concurrency never visits the batch sizes where speculation flips from win to loss — you need a Poisson arrival process to see the crossover at all.

Failure modesWhen it goes wrong

  • Excellent demo, worse production throughputspeculation left enabled at high batch size, where draft tokens compete with real tokens for saturated tensor cores.
  • The speedup disappears on a new workloadacceptance rate collapsed; usually n-gram drafting applied to output that does not copy the input.
  • Out-of-memory or reduced maximum concurrency after enabling a draft modeldraft weights took memory that was previously KV cache blocks.
  • Stuttering token stream in the UIbursty delivery of accepted runs; the server is fine, the client's per-token animation is not.
  • Regression tests fail byte-exactly with identical seedsthe verify pass takes a different arithmetic path than plain decode and flips near-ties.

PapersWhere this comes from

  • Fast Inference from Transformers via Speculative DecodingLeviathan, Kalman & Matias, 2023 (arXiv:2211.17192). Introduced draft-and-verify with an acceptance rule that provably preserves the target model's distribution, and derived the expected-acceptance formula this page uses. It is the reason speculation is an optimisation rather than a quality trade.
  • Accelerating Large Language Model Decoding with Speculative SamplingChen et al., 2023 (arXiv:2302.01318). The concurrent DeepMind formulation, demonstrated at Chinchilla scale, establishing that the technique holds at production model sizes rather than only on toy pairs.
  • Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsCai et al., 2024. Replaced the separate draft model with extra heads on the target model, removing the memory cost, and introduced typical acceptance — the relaxed criterion that trades exactness for a higher acceptance rate.
  • Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative DecodingXia et al., 2024. Surveys drafting strategies and, more usefully for an operator, catalogues the conditions under which reported speedups fail to materialise.
A living map of modern AI — kept current every morning