Home › Deployment, Inference & LLMOps › Continuous batching
⚙️ · Operate

Continuous batching

Continuous batching: iteration-level scheduling that keeps LLM serving GPUs busy across concurrent requests.

In one line

Continuous batching admits and evicts requests at every token step instead of per whole batch, so a serving GPU stays saturated even when responses vary wildly in length.

ConceptWhat it is

Continuous batching (also called iteration-level scheduling, or in NVIDIA's terms in-flight batching) is an LLM inference-serving optimization. The problem it solves comes from how transformers generate: autoregressive decoding emits one token at a time, and different requests finish after wildly different numbers of tokens. Naive static batching gathers a fixed group of requests, runs them together, and cannot free a slot until the whole batch is done, so a batch carrying one 800-token answer keeps its slot occupied while shorter answers sit finished, wasting GPU cycles. This is head-of-line blocking.

Continuous batching moves the scheduling decision down to the granularity of a single decode step. After every forward pass, the scheduler can evict sequences that just hit a stop token or their length cap and admit newly arrived requests into the running batch. The batch composition changes continuously, so the GPU processes a full, freshly-topped-up batch on nearly every iteration. The idea was formalized in the Orca paper and popularized in production by vLLM.

How it worksThe mechanics

Requests land in a queue. The scheduler assembles a running batch subject to a KV-cache memory budget, then the engine runs one forward pass that produces exactly one new token for every active sequence. Immediately after that step the scheduler inspects each sequence: any that emitted an end-of-sequence token or reached its max-length is removed and its KV cache is freed; the freed capacity lets waiting requests be admitted, which are prefilled and joined into the batch. The loop then repeats for the next token step, so eviction and admission interleave continuously rather than waiting for a batch boundary. Newly admitted prompts undergo a compute-heavy prefill phase, which some engines split with chunked prefill so it does not stall the ongoing decode steps.

At a glanceSee it

Continuous batching diagram
Continuous batching diagram 1

Why static batching wastes the GPU — a sequence that finishes early leaves its slot idle until the longest sequence in the batch is done, so throughput is capped by the tail.

Continuous batching diagram 2

The two-phase reality inside each admitted request — a compute-bound prefill of the whole prompt must merge into the memory-bound decode stream, and chunked prefill keeps that merge from stalling the sequences already decoding.

When to use itWhere it fits

  • Any multi-user or multi-tenant serving where many requests arrive concurrently and you want to maximize tokens per second per GPU.
  • Workloads with highly variable output lengths, such as chat, agents, and tool-calling loops, where static batching wastes the most GPU time.
  • Cost-sensitive deployments aiming to serve the same traffic on fewer accelerators by lifting utilization.
  • Bursty, unpredictable traffic where a fixed batch window would either add latency or run half-empty.

When NOT to use itLimits & anti-patterns

  • Single-stream, batch-size-one offline generation, where there is nothing to batch and the feature adds no benefit.
  • Ultra-strict per-request determinism or latency-jitter budgets, since batch composition varies from step to step.
  • Tiny models on CPU or where the GPU is not the bottleneck, so utilization gains are marginal.
  • Environments locked to an engine or runtime that does not implement it, since it cannot be bolted on externally.

Trade-offsAdvantages & costs

Advantages
  • Much higher GPU utilization and throughput than static batching, often several-fold on mixed-length traffic.
  • Lower latency for short requests, which no longer wait behind the longest response in a fixed batch.
  • Better cost per token, translating directly into fewer GPUs for the same load.
  • Naturally adapts to real-world bursty, variable-length request streams without a tuned batch-wait window.
Trade-offs & costs
  • The serving engine must support it natively; it is a scheduler-level feature, not a wrapper you add.
  • KV-cache memory pressure can force preemption, recomputation, or swapping when too many long sequences run at once.
  • Prefill of newly admitted requests can interfere with ongoing decode steps, causing latency bubbles unless chunked prefill is used.
  • Variable batch composition makes tail latency, fairness, and SLA capacity planning harder to reason about.

ExampleIn the real world

A team serves a chat assistant to many concurrent users on a single high-end data-center GPU. With static batching of 32, any batch that includes one long 800-token reply holds all 32 slots until that reply finishes, so the GPU runs mostly empty as short replies complete early and wait. Switching to a continuous-batching engine, finished replies are evicted at each token step and queued requests immediately fill the freed slots, so the batch stays near-full every iteration. The result is a several-fold throughput increase and a lower median time-to-completion, letting the team absorb the same traffic on fewer GPUs. The main new operational task is tuning the KV-cache memory fraction and max batch size so that peak concurrency does not trigger frequent preemption.

ToolsHow to implement it

  • vLLMcontinuous batching paired with PagedAttention for efficient KV-cache memory management.
  • NVIDIA TensorRT-LLMimplements the same idea under the name in-flight batching.
  • Hugging Face Text Generation Inference (TGI)production server with continuous batching built in.
  • SGLangand LMDeploy — high-throughput serving engines that schedule at the iteration level.

Cost & effortWhat it takes

Continuous batching is a feature of the serving engine rather than something you build, so the adoption cost is mostly configuration and validation, not engineering. In the mainstream open engines it is on by default, so the real effort goes into tuning knobs, the KV-cache memory fraction, maximum batch size and sequence count, and chunked prefill, then load-testing to confirm the throughput gain holds without violating latency SLAs under peak concurrency. The payoff is a direct reduction in GPU count and cost per token, which usually repays the tuning effort quickly for any busy multi-user deployment.

A living map of modern AI — kept current every morning