Let a cheap draft model propose tokens and the expensive target model verify them in parallel, cutting decoding latency two to three times with no change to the output.
ConceptWhat it is
Speculative decoding is an inference-time optimization that speeds up text generation without changing what the model outputs. A large, slow target model normally emits one token per forward pass, and because each pass is memory-bandwidth bound — dominated by streaming the weights through the GPU rather than by arithmetic — a single token barely uses the hardware. Speculative decoding fills that idle capacity by letting a small, cheap draft model guess several tokens ahead, then having the target model check the whole guess in one shot.
The defining property is that it is lossless: a verification step accepts only the prefix of guessed tokens the target would itself have produced, so the output distribution is preserved exactly. You pay for extra draft compute and a second model to run, and in return you collapse many sequential steps into one, typically 2–3× faster on latency-sensitive workloads.
How it worksThe mechanics
At each step the draft model autoregressively generates a short block of K candidate tokens, which is cheap because the draft is small. The target model then runs a single forward pass over the whole block in parallel, producing its own probability at every position. Walking left to right, each candidate is accepted while it agrees with the target's choice; at the first disagreement the candidate is rejected and the target's own token is substituted, which guarantees correctness. Every accepted token plus the one correction advance the sequence, so a well-aligned draft yields several tokens for the cost of a single target pass, and the loop repeats until the response is complete.
At a glanceSee it
Zooming into one token, the probabilistic accept-or-resample rule is exactly what keeps the sped-up output identical to plain target sampling.
A taxonomy of draft sources — a separate small model, the target self-drafting via extra heads or hidden features, or a model-free retrieval lookup.
When to use itWhere it fits
- Latency-sensitive interactive generation — chat, coding assistants, agents — where time-to-token is felt directly.
- Serving a large target model that is memory-bandwidth bound at small batch sizes, with GPU compute to spare.
- Predictable, low-entropy text such as code, structured output, or boilerplate, where the draft guesses well and acceptance runs high.
- When you need the exact outputs of the large model, ruling out lossy shortcuts like aggressive quantization.
When NOT to use itLimits & anti-patterns
- No well-aligned draft model exists and building or finding one is not worth the effort — acceptance stays low and gains evaporate.
- Throughput-bound batch serving where large batches already saturate the GPU, leaving no idle capacity to reclaim.
- Highly creative, high-entropy generation where the draft rarely agrees with the target, so most tokens are rejected.
- Very small target models, where the draft's overhead cancels any meaningful savings.
Trade-offsAdvantages & costs
Advantages
- 2–3× lower decoding latency on suitable workloads with no change to output quality.
- Lossless — it preserves the target model's exact output distribution rather than approximating it.
- Drops into existing serving stacks; several production frameworks support it out of the box.
- Composes cleanly with other optimizations like quantization and KV caching.
Trade-offs & costs
- Requires a well-aligned draft model; a mismatched one delivers little or even negative speedup.
- Adds GPU memory and orchestration cost for running two models in lockstep.
- Speedup is workload-dependent and hard to predict, since acceptance rate varies prompt to prompt.
- Draft compute spent on rejected tokens is wasted, which caps the best-case gain.
ExampleIn the real world
Consider a team serving a 70-billion-parameter chat model behind a coding assistant, where users feel every millisecond of streaming latency. They pair it with a same-family 1-billion-parameter model as the draft: the small model proposes eight tokens at a time, and the large model verifies the block in a single forward pass. On the assistant's typical output — code, closing brackets, repetitive scaffolding — the draft is right most of the time, so accepted blocks are long and end-to-end token throughput roughly doubles, while the generated code is byte-for-byte what the 70B model would have produced on its own.
ToolsHow to implement it
- vLLMproduction serving engine with built-in speculative decoding via a draft model, n-gram matching, or prediction heads.
- TensorRT-LLMNVIDIA's inference library with speculative decoding and Medusa/EAGLE support.
- Hugging Face transformersassisted generation, the reference implementation of draft-model speculative decoding.
- Medusa / EAGLEself-drafting methods that add lightweight prediction heads to the target model instead of running a separate draft.
Cost & effortWhat it takes
The engineering lift is modest when a suitable draft already exists — often a smaller checkpoint from the same family — and the serving framework supports it: mostly configuration plus tuning the speculation length K and confirming quality is unchanged. The recurring cost is the extra GPU memory to hold the draft and the draft compute wasted on rejected tokens; the payoff is 2–3× lower latency, which for interactive products means better responsiveness and, at fixed hardware, more requests served per GPU. If no aligned draft exists, self-drafting approaches like Medusa need a short training run to fit the prediction heads.