Temperature reshapes the whole distribution before truncation runs, so top_p and top_k mean completely different things at 0.2 than at 1.2 — tune one, not both.
Why you'd careThe thing you have already noticed
Your extraction prompt works on four hundred test inputs and then, in production, returns a field name that is not in your schema. Nothing changed. You are running temperature=0.7 with top_p=0.95 because those were the values in the example you copied. That combination leaves a tail of tokens holding a fraction of a percent of the mass, and over a hundred thousand calls generating three hundred tokens each you draw from that tail tens of thousands of times. Sampling is the only stage in the entire pipeline that deliberately introduces randomness. Everything upstream is deterministic in intent; this is where the intent gets spent, and the two numbers that spend it are almost always set by copy-paste.
In and outWhat goes in, what comes out
| In | A logits vector of length vocab_size for the current position, already modified by penalties and by any constraint mask that set illegal entries to −inf. Plus the per-request sampling parameters and, if a seed was supplied, that sequence's own RNG state. |
|---|---|
| Process | Divide every logit by temperature, softmax to probabilities, then truncate by one or more rules: keep the top_k highest; keep the smallest set whose cumulative mass reaches top_p; keep everything above min_p × p_max. Renormalise the survivors and draw one index with a multinomial sample. |
| Out | A single token id, appended to the sequence and to the KV cache on the next step. Optionally the logprob of the chosen token and the top-n alternatives. The RNG state advances by one draw. |
The full distribution is destroyed here and nothing downstream can recover it. Unless you explicitly requested logprobs, the only surviving record that the second choice sat at 0.49 against 0.51 is the token you got. This is why an answer that “went wrong at token 40” cannot be diagnosed after the fact, and why re-running it under different batch conditions may not reproduce it. Sampling is also the one irreversible commitment in autoregressive decoding: the token joins the context and conditions everything after it.
ConceptThe idea underneath
Temperature is a division inside the softmax: p_i = exp(z_i / T) / sum_j exp(z_j / T). Because it divides the logits, it scales the gaps between them, not the values. As T approaches 0 the largest gap dominates and you get argmax; as T grows the gaps vanish and you approach uniform. It is multiplicative in log space, which is why it is not interchangeable with anything additive.
Truncation exists because a maximum-likelihood-trained model assigns small but nonzero probability to tens of thousands of implausible tokens, and the total of that tail is not small. Holtzman's observation was that if you sample from it often enough, you eventually emit a token the model itself considers unlikely — and then you condition on it. The model has never seen text like that, so its next distribution degrades, and the errors compound. Truncation is a guardrail against falling out of the training distribution.
The three rules differ in what they are sensitive to. top_k keeps a fixed count and ignores shape entirely: k=50 admits junk when the distribution is genuinely peaked and deletes real options when it is genuinely flat. top_p adapts the count to the mass, but it is thrown by a single dominant token — if the best token holds 0.94 and top_p is 0.95, you sweep in a very long tail just to reach the threshold. min_p is a relative threshold: keep p_i ≥ min_p × p_max. That makes it scale-free with respect to the model's confidence, cutting hard when the model is sure and staying wide when it is not. Hewitt et al. formalised this family as undoing a smoothing artefact of MLE training.
One implementation detail that bites: the order of operations differs between engines. Check whether temperature is applied before or after truncation in yours — the same two numbers mean different things either way.
At a glanceSee it
Logits become probabilities, get truncated by one of three rules, then one token is drawn.
The knobsHyperparameters and nuance
- temperature0 to 2, default 1.0 on OpenAI-compatible APIs. Most servers special-case 0 to greedy rather than dividing by zero. Above roughly 1.3, JSON structure and grammar start to break; at 0, long-form generation falls into verbatim loops.
- top_p0 to 1, default 1.0; 0.9–0.95 is the common range. Below about 0.5 the output goes bland and repetitive. At 1.0 the entire tail is live and rare-token accidents happen at whatever rate the model's tail mass implies.
- top_kinteger; disabled by default on hosted APIs (encoded as 0 in current vLLM, −1 in older releases), 40–64 typical in local stacks, with llama.cpp defaulting to 40. Small
kon a legitimately flat distribution deletes correct alternatives; largekdoes nothing thattop_pwas not already doing. - min_p0 to 1, useful range 0.02–0.1. The default differs by engine and this catches people out: vLLM defaults to 0 (off), while llama.cpp's server defaults to
0.05, so min-p is already shaping your output on a stock llama.cpp before you set anything. Not exposed by OpenAI's or Anthropic's public parameter sets. Above about 0.2 it collapses to near-greedy. - seedinteger. Makes the sampler's draw reproducible. It does not make the logits reproducible, which is the reason seeded requests still vary on a shared server.
- nnumber of samples drawn from one prefill.
n>1amortises the prompt but multiplies decode cost and, with a shared seed, the samples are correlated in ways people rarely check. Do not reach forbest_ofas its companion: on OpenAI it exists only on the legacy Completions API, never on Chat Completions, and vLLM removed it when the V1 engine replaced V0.
EffectHow this stage moves the answer
Raise temperature and the first things to go are proper nouns, dates and numbers — precisely the positions where the correct answer is a single peaked token that only temperature can dislodge. Next go JSON keys and enum values, which look like free choices to the sampler and like a contract to you. Last and most visibly, the model starts inventing plausible citations and API methods, because a fabricated name is a perfectly ordinary sample from a flattened distribution. Drop to temperature 0 and the failure inverts: long-form generation loops, lists repeat items, and the model restates the same sentence with small variations until it hits max_tokens. Worth internalising: greedy decoding is not “the best answer”. It picks the highest-probability token at each step, which is a different thing from the highest-probability sequence, and on long outputs the two diverge badly.
EvalsWhat it does to your measurements
This is the most under-reported variable in published evaluation numbers. The same model at temperature 0 and at 0.7 can differ by several points on GSM8K or HumanEval, and the sampling config is frequently absent from the reported setup. pass@k is only defined relative to a temperature — the original Codex evaluation used a low temperature for pass@1 and a higher one for larger k, because the metric rewards diversity as k grows. Comparing someone's pass@1 at 0.2 to your pass@1 at 0.8 is not a comparison. The other trap is variance: a single greedy run reported as the score hides the fact that the model's answer at temperature 0.7 has a distribution, and the mean of that distribution is what your users experience. Report n, report the spread, and log temperature, top_p, top_k and min_p with every number.
Failure modesWhen it goes wrong
- Valid JSON 99.7% of the time and invalid the resttail sampling at
top_p=0.95; either lower it, or make it structurally impossible with constrained decoding. - The model repeats a paragraph until it hits the token limitgreedy or near-greedy decoding on long-form text with no penalty configured.
- Quality changes after a server upgrade with no config changethe engine changed a default (
top_kfrom disabled to 50 is the classic) or changed the order in which temperature and truncation are applied. - A reasoning model produces incoherent chainsit was run at the sampling settings tuned for a non-reasoning model; several vendors require a fixed temperature with thinking enabled and ignore or reject the rest.
- Seed is set and the output still variesthe seed governs the draw, not the logits; see determinism and batch invariance.
PapersWhere this comes from
- The Curious Case of Neural Text DegenerationHoltzman et al., 2019 (arXiv:1904.09751; ICLR 2020). Introduced nucleus (top-p) sampling and demonstrated that likelihood-maximising decoding produces degenerate, repetitive text while pure sampling wanders off-distribution. Every truncation parameter on your API exists because of this result.
- Hierarchical Neural Story GenerationFan, Lewis & Dauphin, 2018 (arXiv:1805.04833; ACL 2018). Introduced top-k sampling for open-ended generation; it is the ancestor of the
top_kparameter and the reason a fixed cutoff was the first thing anyone tried. - Truncation Sampling as Language Model DesmoothingHewitt, Manning & Liang, 2022 (arXiv:2210.15191; Findings of EMNLP 2022). Reframed truncation as removing a smoothing artefact introduced by maximum-likelihood training, and showed concretely that top-p over-truncates at confident positions. Its own proposals are eta-sampling and epsilon-sampling, whose threshold is entropy-dependent rather than scaled to the top token — read it as the theoretical case for adaptive truncation in general, not as the origin of min-p.
- Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM OutputsNguyen et al., 2024 (arXiv:2407.01082; ICLR 2025 oral). The actual min-p paper: the threshold is the top token's probability times
min_p, so the cutoff tightens automatically where the model is confident. Note the order of events — min-p shipped as an engine feature in llama.cpp and vLLM well before this was published, and the write-up has itself drawn a published critique, so treat the reported gains as contested rather than settled.