Several completely different events end a generation, and the only thing that tells them apart is a field most clients never read.
Why you'd careThe thing you have already noticed
You get a JSON parse error from a call that worked all week, and it will not reproduce on the same prompt locally. Or a list that should have ten items has seven, and the seventh looks complete. Or a retry loop runs the same request five times and collects five refusals. All three are the same omission: the response carried a field saying exactly why generation ended, and the code treated a length cap, a matched delimiter, a clean finish and a refusal as one event. In a chat window those four look nearly identical — often the only visual difference is a missing closing brace — which is precisely why the field exists and precisely why ignoring it fails silently.
In and outWhat goes in, what comes out
| In | The decoded token stream one token at a time, plus every configured terminator: max_tokens, the stop-sequence list, the EOS and end-of-turn token IDs baked into the tokenizer config, a server-side deadline, and any guard verdict able to pre-empt all of them. |
|---|---|
| Process | After each token the server tests the terminators in a fixed, implementation-defined precedence and the first match wins. Stop-string matching runs on the incrementally detokenized suffix rather than on token IDs, so it must cope with a stop string that straddles a token boundary. |
| Out | The final text plus a typed reason: end_turn, max_tokens, stop_sequence or tool_use on the Anthropic Messages API; stop, length, tool_calls or content_filter on OpenAI. When a stop sequence matched, the matched string is reported separately and by default excluded from the text. |
The text survives intact; the reason is compressed to a single label for the whole completion, so you learn that generation was capped but not which of ten list items was the last complete one. The KV cache for the sequence is released at this point, so continuing is a fresh request that must re-prefill the whole prompt unless prefix caching happens to catch it.
ConceptThe idea underneath
Two mechanisms with nothing in common hide behind one field. The first is a learned symbol. Every model has a vocabulary entry for end-of-sequence — in chat models usually an end-of-turn token — and at each step the network produces a distribution over the entire vocabulary with that token in it. Generation ends when it is sampled. So the model decided to stop is a probabilistic event with a readable probability, P(EOS | tokens 1..t), and at any temperature above zero the same prompt can legitimately end at different lengths on different runs. That probability is calibrated entirely by the length distribution of the post-training data, which is why a model fine-tuned on terse assistant replies truncates a long structured task while one trained on verbose answers runs to the cap.
The second mechanism is not learned at all. Stop sequences are a string condition applied to a token stream by the server, after incremental detokenization. The model never sees them and cannot cooperate with them. Because token boundaries and string boundaries do not line up, a stop string can be produced as part of a longer token, and the server then has to decide how much of that token to keep — the source of both over-trimmed output and stop strings that mysteriously fail to match at all.
Everything else is precedence. max_tokens is a hard counter on output tokens; the deadline is wall-clock; a guard verdict pre-empts both. When several conditions become true on the same step, which one gets reported is an implementation choice rather than a specification. Treat the reason enum as open and always keep a default branch: providers add values over time, and a runtime cast into a closed enum will throw on the first new one it meets in production.
At a glanceSee it
Four different terminators collapse into one typed reason the client must branch on.
The knobsHyperparameters and nuance
- max_tokens
max_completion_tokenson current OpenAI models. Caps output tokens, not characters, and on reasoning models the hidden reasoning counts against it. Set too low it truncates mid-structure; set too high it lets a looping generation burn the entire budget before you find out. - stop / stop_sequencesup to four strings on OpenAI, matched against the detokenized suffix. A stop string that can occur inside legitimate output, such as a double newline, truncates a subset of your traffic and gives no other signal that it did.
- include_stop_str_in_outputvLLM, default false. Flipping it changes the contract for every downstream parser at once; it is the setting that makes a delimiter appear or vanish with no model change behind it.
- ignore_eosvLLM, default false. Forces decoding past the EOS token. It exists for throughput benchmarking; leave it enabled in production and every single request runs to max_tokens.
- min_tokensvLLM. Suppresses EOS for the first N output tokens. It is the right fix for a model that stops after one sentence; set too high it manufactures padding and hedging the model did not want to write.
- request timeout / server deadlinethe terminator with the worst reporting. A deadline hit usually surfaces as a transport error rather than as a stop reason, so it is the one truncation your stop-reason branch will never see.
EffectHow this stage moves the answer
Nothing here changes the tokens the model produced; it changes whether your application notices what happened to them. Ignore the field and a max_tokens truncation reaches the user as an answer that simply stops — a list ending at item seven, a sentence with no verb, a code block with no closing brace — and it reads as the model being incompetent rather than as a configuration bug. Ignore it in a tool loop and a refusal gets retried as though it were a network error, so the same request runs five times and produces five refusals and five bills. Handle it and each terminator becomes a different repair: on max_tokens you can prefill the partial assistant turn and ask for the remainder; on stop_sequence you know the content is complete and your delimiter did its job; on tool_use you dispatch rather than render; on a filter reason you surface it instead of retrying.
EvalsWhat it does to your measurements
This is where a harness quietly scores a model as worse than it is. A completion cut off at max_tokens fails exact match, fails a unit test, and gets marked wrong by an LLM judge, so a max_tokens default 200 tokens too low is indistinguishable in the aggregate from a genuine model regression. Reasoning models make it acute: a harness defaulting to 512 output tokens will spend all of them on hidden reasoning and score near zero on everything, with no error anywhere in the logs. The discipline is to record a stop-reason histogram beside every score and treat any run with more than a percent or two of length stops as invalid rather than as a result. Stop sequences fail more subtly still — a newline-based stop string truncates only the multi-paragraph answers, biasing the score toward short outputs in a way no aggregate number will ever reveal.
Failure modesWhen it goes wrong
- JSON parse errors in production that never reproduce locallylonger real-world inputs push the completion into the max_tokens cap; the stop reason says so and nothing in the client reads it.
- Every long answer stops at almost exactly the same placethe cap counts output tokens, so it moves with tokenizer efficiency, and non-English or code-heavy answers hit it far sooner than the character count suggests.
- A retry loop hammers the API and returns the same thing five timesa refusal or filter reason routed through the generic transient-error path.
- Downstream parsers break because the delimiter disappearedstop sequences are excluded from the output by default, so the text ends just before the marker the parser is looking for.
- The model answers in one sentence a question that needs tenEOS probability mass inherited from short fine-tuning data; min_tokens or an explicit length instruction is the lever, not temperature.
PapersWhere this comes from
There is no research literature for this stage, and reaching for some would be dishonest: it is an API contract, not a modelling problem. The normative sources are the provider references for the terminal-reason field — stop_reason and stop_sequence in the Anthropic Messages API, finish_reason in OpenAI's Chat Completions — plus the WHATWG server-sent events specification for how the terminal frame reaches the client. For actual behaviour, read implementations: HuggingFace transformers exposes StoppingCriteria and GenerationConfig, and vLLM's SamplingParams gathers the whole terminator set in one place, including stop_token_ids, include_stop_str_in_output, ignore_eos and min_tokens. The one genuinely adjacent research result comes from machine translation rather than serving: Stahlberg and Byrne, 2019, showed that exact search in neural MT models very often prefers the empty string, which is direct evidence that a model's probability mass on end-of-sequence is a modelling artifact and not a trustworthy signal about how long an answer should be.