Temperature zero fixes the sampler, not the logits, and the arithmetic that produced those logits depends on who else happened to be in your batch.
Why you'd careThe thing you have already noticed
You wrote a regression test: fixed prompt, temperature=0, seed=42, assert the exact output string. It passed for two weeks and then failed in CI on a Tuesday afternoon with a one-word difference. You reran it and it passed. Nothing had been deployed. This is not flakiness in your harness. Setting temperature to zero makes the sampler deterministic — it takes the argmax — but it does nothing about how the logits were computed. On a shared server the reduction order inside the attention and matmul kernels depends on how many other sequences were batched alongside yours, and floating-point addition is not associative, so the same prompt yields logits that differ in the last few bits. When the top two tokens are close, that flips the argmax.
In and outWhat goes in, what comes out
| In | A request carrying a seed and sampling config, arriving at a server that will batch it with whatever else is in flight, on a specific kernel build, GPU model, tensor-parallel degree, quantisation scheme and prefix-cache state. |
|---|---|
| Process | Kernels reduce over the hidden and sequence dimensions using a split strategy selected at runtime from the tensor shapes. Split-K matmuls, atomic accumulation and FlashAttention's log-sum-exp rescaling all combine partial sums in an order that depends on batch composition and tile mapping. |
| Out | Logits that agree to roughly fp32 or bf16 precision but not bit for bit across runs, and therefore an argmax that agrees almost always and occasionally does not. A generation that is semantically stable and not byte-stable. |
The distribution is preserved to within numerical tolerance. Bitwise reproducibility is lost, and lost permanently: you cannot recover which batch a past request ran in, so a production output can never be reproduced exactly after the fact unless the server pinned a batch-invariant kernel path and recorded the build. Divergence also compounds — one flipped token conditions every token after it, so a last-bit difference becomes a visibly different paragraph.
ConceptThe idea underneath
There are two independent sources of nondeterminism here and they are constantly conflated.
The first is the sampler. A seed, a per-request RNG stream, and greedy decoding solve it completely. Most teams stop here and are then baffled when the problem persists.
The second is the kernels, and it is arithmetic rather than randomness. IEEE 754 addition rounds after every operation, so (a + b) + c is not equal to a + (b + c) in general. GPU reductions split a sum across threads and blocks and combine the partial results in whatever order the tile decomposition produces — and that decomposition is chosen from the tensor shapes, of which the batch dimension is one. So the value computed for your row depends on how many other rows were present. A kernel is batch-invariant when it does not; almost no fast inference kernel is, because the fast path deliberately picks different split-K factors and tile sizes at different batch sizes.
The Thinking Machines Lab post argued that this, rather than nondeterministic atomics, is the dominant cause of LLM inference nondeterminism in practice: most serving kernels already avoid atomics, and yet output still varies, because none of them are batch-invariant.
Why it surfaces at all is a property of language. Next-token distributions are frequently very flat — the token after “the”, a whitespace choice, a synonym — and at those positions the top two logits can sit within 1e-4 of each other. A relative perturbation of 1e-3 in bf16 is larger than that gap. Most positions are safe. You only need one. Chunked prefill, prefix caching and continuous batching all change which tokens are processed in which chunk, so they move the reduction boundaries too.
At a glanceSee it
The same prompt takes a different reduction path depending on who else is in the batch.
The knobsHyperparameters and nuance
- seedper-request integer. Makes the sampler's draw reproducible and nothing else. On OpenAI-compatible APIs it is documented as best-effort precisely because the logits underneath are not guaranteed stable.
- temperature = 0removes sampler randomness, but it also removes all averaging, so it amplifies sensitivity to logit ties. Counter-intuitively, a seeded run at temperature 0.7 and a greedy run are both nondeterministic, and the greedy one fails more visibly.
- system_fingerprint(OpenAI-compatible) — a backend build identifier returned with the response. If it differs between two runs, reproducibility was never available and any comparison you drew is void. Log it with every eval result.
- batch-invariant kernel modevLLM has shipped an opt-in batch-invariant execution path following the Thinking Machines work. It buys bitwise stability across batch sizes at meaningful throughput cost. Implementation-specific, evolving, and worth verifying against your own build rather than a release note.
- --enable-prefix-caching(vLLM) — a cache hit changes which tokens get recomputed and therefore the reduction structure, so identical requests can produce different bytes depending on cache state. Disable it when you are chasing reproducibility, and expect a large latency regression while you do.
- torch.use_deterministic_algorithms(True) and CUBLAS_WORKSPACE_CONFIGmakes individual PyTorch operations reproducible run to run at a fixed shape. It does not give batch invariance, which is why people try it, see no improvement, and conclude the problem is elsewhere.
EffectHow this stage moves the answer
Semantically, usually nothing. Two runs that diverge at a low-information token typically reconverge on meaning within a sentence; the answer is correct both times, phrased differently. The damage lands downstream of the answer, on anything that treats the output as a key: exact-match assertions, response caches, audit-trail hashes, golden-file tests over generated code, snapshot tests of tool-call arguments. The case that actually hurts is when divergence lands on a high-information token — the first digit of a figure, the choice between two API method names, a yes-or-no classification sitting at the decision boundary. There the answer genuinely changes, at an unpredictable and low rate that no amount of retrying reveals. A classifier that is 99.8% self-consistent is a perfectly good product and an infuriating regression suite, and confusing those two is how teams end up disabling tests instead of fixing tolerances.
EvalsWhat it does to your measurements
This is why the same model scores differently on the same benchmark twice. Sensitivity concentrates in exact-match and generation-based multiple choice; rubric-graded free-form metrics barely move. Three concrete ways it invalidates a run. A benchmark executed at concurrency 32 and rerun at concurrency 1 is not the same experiment, because the kernel paths differ. A cache-warmed run is not comparable to a cold one — prefix caching changes the numerics, not only the latency, so “warm up the cache first” silently changes the measurement. And comparing two models on one harness at different effective batch sizes attributes to the models a difference that is partly arithmetic. The practice that works: at least three samples per item with a tolerance band, system_fingerprint or server build logged with every score, concurrency recorded, and no single greedy run ever treated as ground truth.
Failure modesWhen it goes wrong
- A CI regression test flakes on one token and passes on rerunbatch composition changed the reduction order in the kernels; the model and the seed are identical.
- Local single-request output differs from production outputdifferent batch size, and often a different tensor-parallel degree, selecting a different kernel path.
- The same seed gives a different answer after a server upgradekernel selection, fused-operation changes or a new attention backend; the seed only ever governed the sampler.
- Answers change when prefix caching is turned ona cached prefix skips recomputation, so the surviving arithmetic path is not the one that produced the original.
- Two GPUs of the same model disagreediffering SM counts change tile mapping and therefore reduction order; the same applies across any change in tensor-parallel degree.
PapersWhere this comes from
- Defeating Nondeterminism in LLM InferenceThinking Machines Lab, 2025. An engineering blog post, not a peer-reviewed paper, and worth saying so. It argued that lack of batch invariance rather than nondeterministic atomics is the dominant cause in practice, and demonstrated batch-invariant kernels producing bitwise-identical output across batch sizes. It is the clearest published account of the mechanism.
- IEEE Standard for Floating-Point Arithmetic (IEEE 754-2019)IEEE, 2019. The normative source for why addition does not associate under rounding. Everything on this page is a consequence of a specification, not of a bug.
- PyTorch reproducibility documentation and torch.use_deterministic_algorithmsPyTorch project, ongoing. The reference implementation of per-operation determinism, and the place most teams learn the wrong lesson: it fixes op-level reproducibility at fixed shape and does nothing about batch invariance. Beyond these, there is essentially no peer-reviewed literature on serving determinism; the material is blog posts, vendor docs and kernel source.