Home › What Happens After You Hit Enter › Continuous (in-flight) batching
Pipeline stage · Operate

Continuous (in-flight) batching

Add and remove requests from the running batch at every token step, instead of waiting for the whole batch to finish.

In one line

The batch is rebuilt at every token, so your latency depends on who else is in it and the same prompt at temperature 0 can produce different text under different load.

Why you'd careThe thing you have already noticed

The same prompt, same model, same parameters returns in 900 milliseconds at 2am and 4 seconds at midday, and no log shows a retry or an error. Or you benchmarked locally one request at a time, got a number you liked, and watched it evaporate in production. Continuous batching is why. The GPU does not process your request; it processes a step, for everyone currently generating, all at once. At every token boundary the scheduler rebuilds that group — finished sequences leave, queued ones join, and requests under memory pressure can be evicted and restarted. Your throughput is a property of the group you landed in. This is also why serving throughput improved by roughly an order of magnitude over naive batching with no change to the model.

In and outWhat goes in, what comes out

InA queue of pending requests, each with a prompt, sampling parameters and a token budget; the set of currently generating sequences with their cache block tables; and the free-block count of the KV pool.
ProcessEach step the scheduler selects sequences fitting the token budget and available blocks, gathers their next-token inputs into one batch, runs a single forward pass, appends one token per sequence, frees the blocks of anything that finished, and admits waiting requests immediately.
OutOne token per running sequence, streamed to its own connection; an updated block pool; and a scheduling decision about what runs on the next step.

Logical isolation is preserved — no request can read another's cache and each carries its own sampler state. Numerical isolation is not. The batch's shape determines tile sizes and reduction order inside the kernels, so a request's logits depend on the shape of the batch it happened to ride in. Bit-exact reproducibility across runs is the thing given up here.

ConceptThe idea underneath

There is no machine learning in this stage, and it is worth saying so plainly. It is a scheduling result, and the enabling property is that a decode step for one sequence is independent of every other sequence given its own KV cache. Sequences can therefore be gathered into a batch arbitrarily and regrouped every step, because nothing carries between steps except cache belonging to a specific sequence.

Why it matters is arithmetic intensity again. A decode step reads the entire weight matrix out of HBM in order to multiply it by one token's activations — at batch size 1 that is roughly two FLOPs per byte read, on hardware that wants hundreds. Batching 64 sequences reads the same weights once and does 64 times the arithmetic, so tokens per second per GPU rises almost linearly with batch size until either compute saturates or the KV cache runs out of room. Batching converts a bandwidth-bound operation into a compute-bound one.

Static batching wastes this because completion times vary enormously. Group 32 requests and, when 31 have finished at 40 tokens, the batch still runs for the one writing an essay with 31 idle slots. Orca's contribution was iteration-level scheduling: make admission and eviction decisions every token rather than every request. Combined with paged KV allocation, which lets a freed sequence's memory be reused immediately without fragmentation, that is what produced the order-of-magnitude throughput jump.

The cost is coupling. Every request now shares a step with strangers, which is why latency percentiles rather than averages are the only honest way to describe a deployment.

At a glanceSee it

Continuous (in-flight) batching diagram

The scheduler rebuilds the batch every token step, admitting and retiring sequences continuously.

The knobsHyperparameters and nuance

  • max_num_seqsthe cap on concurrent sequences in a batch. The V1 engine defaults to 1024, up from 256 in V0, which is itself a common cause of unexpected out-of-memory on upgrade. Higher raises throughput and per-token latency for everyone; lower gives crisp latency and leaves the GPU idle between steps.
  • max_num_batched_tokenstotal tokens per step, prefill and decode combined. Once chunked prefill is on, this is the real throughput-versus-latency dial rather than the sequence count.
  • gpu_memory_utilizationthe fraction of VRAM the server claims, default 0.90. Whatever is left after weights becomes KV cache and therefore sets concurrency. Push it to 0.98 and activation spikes become out-of-memory crashes under load.
  • scheduling policyfirst-come-first-served by default; priority scheduling lets short interactive requests jump long batch jobs. Without it, a single 8,000-token generation sets tail latency for everything behind it.
  • preemptionunder KV pressure vLLM evicts sequences. V0 offered preemption_mode as a choice between recompute, which discards and re-prefills them, and swap, which moves their blocks to host memory. V1 implements recompute only and does not accept --preemption-mode, so on current vLLM this is a behaviour to anticipate rather than a dial to turn: frequent preemption is the signal to raise gpu_memory_utilization or lower max_num_seqs.
  • max_model_lenindirectly a concurrency knob, because reserving room for the worst-case context reduces how many sequences fit at once.

EffectHow this stage moves the answer

The claim that this stage causes no quality change is right about expected quality and wrong about determinism, and that difference matters when you are debugging. Because reduction order inside matrix multiplications and attention kernels depends on batch shape, the logits your request receives are a function of who else was in the batch. Most of the time the difference sits in the last few decimal places and the same token still wins. Occasionally the top two candidates are close enough that it does not, and from that token onward the completions diverge completely. So: temperature 0, fixed seed, identical prompt, two visibly different answers, one written at 2am and one at midday, with nothing broken anywhere. Under KV pressure there is a second path, since a preempted request is re-prefilled and resumed, which is exact in principle and yet another chance to meet a different batch shape in practice.

EvalsWhat it does to your measurements

This is the stage that quietly invalidates eval runs. If your harness runs at concurrency 64 and your baseline was captured at concurrency 1, completions will differ at temperature 0 for reasons unrelated to the change you are testing, and exact-match or diff-based scoring will report a regression that does not exist. Pin concurrency, server version and attention backend before comparing runs, or accept that only statistical differences are meaningful. On the performance side, single-stream benchmarks are close to worthless for capacity planning because they measure the one condition production never has. Measure at a fixed request rate, report p50 and p99 for time-to-first-token and inter-token latency separately, and plot throughput against latency rather than quoting either alone — they sit on one curve and any deployment can slide along it by changing max_num_seqs.

Failure modesWhen it goes wrong

  • p99 latency spikes while p50 stays healthyhead-of-line blocking, usually a long prefill occupying a step ahead of a queue of short requests.
  • Requests stall for seconds and then resumepreemption under KV pressure; in recompute mode the sequence is discarded and re-prefilled from scratch.
  • Throughput plateaus well below what the GPU's FLOPs suggestKV cache memory rather than compute is capping batch size. Check free blocks, not utilisation.
  • Out-of-memory crashes only at peak loadgpu_memory_utilization set too high, leaving no headroom for the moment a long prefill and a full decode batch coincide.
  • Temperature-0 outputs differ between two runs of the same evaldifferent batch shapes producing different reduction orders, not a model change.

PapersWhere this comes from

  • Orca: A Distributed Serving System for Transformer-Based Generative ModelsYu et al., 2022. Introduced iteration-level scheduling, admitting and retiring requests every token rather than every batch, which is the mechanism now called continuous or in-flight batching. Published at OSDI; there is no arXiv preprint.
  • Efficient Memory Management for Large Language Model Serving with PagedAttentionKwon et al., 2023. arXiv:2309.06180. Paged KV allocation, which is what makes freeing and admitting sequences mid-flight practical without fragmentation. The vLLM paper.
  • Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAgrawal et al., 2024. Quantified the stall a long prefill inflicts on ongoing decodes and introduced stall-free scheduling, the reason chunked prefill is now on by default.
  • Defeating Nondeterminism in LLM InferenceHorace He and collaborators, Thinking Machines Lab, 2025. An engineering writeup rather than a paper; it traced temperature-0 nondeterminism to batch-size-dependent reduction order in kernels and showed batch-invariant kernels fix it at a throughput cost.
A living map of modern AI — kept current every morning