Stop sequences are matched against decoded text rather than token ids, so the server must hold back output it has already generated in case a stop is half-emitted.
Why you'd careThe thing you have already noticed
Two symptoms, one cause. First: you set a stop on a blank line and the model still returns a blank line in the middle of an answer. Second: your streamed output arrives in slightly lumpy chunks and the final characters sometimes look shaved — an answer ending in Done where you expected Done. Stop sequences are string matches over detokenised text, and the model does not produce characters, it produces tokens. A stop string can be generated by several different token sequences, can be split across two tokens, or can be swallowed inside one longer token that also carries real content. To guarantee it never leaks to you, the server has to withhold any tail that could still become a stop.
In and outWhat goes in, what comes out
| In | A list of stop strings, optionally a list of stop token ids, the token ids generated so far, and the incrementally detokenised text buffer for this sequence. Stop token ids are checked against the sampled id directly, before any detokenisation happens. |
|---|---|
| Process | After each token, append its decoded text to the buffer and scan the tail for any stop string. Withhold the last N-1 characters, where N is the longest configured stop string, since they could still become the prefix of one. Release everything before that boundary as a delta. |
| Out | Either a further text delta released to the client, or a terminated sequence with the stop string truncated from the output unless include_stop_str_in_output is set, plus a finish reason of stop and, on some APIs, which stop string fired. |
The token ids up to the stop survive. What is discarded by default is the stop string itself and, on naive implementations, the whole token that contained it — if the model emitted a single token whose text is a period followed by two newlines and your stop is the double newline, you lose the period too. Also lost is the information that generation would have continued sensibly: finish_reason: stop looks identical whether the model finished its thought or was cut off mid-sentence.
ConceptThe idea underneath
There is no neural network content in this stage at all. It is streaming multi-pattern string matching, complicated by the fact that the alphabet you are matching in — characters — is not the alphabet you are producing in.
The correctness invariant is what generates the cost. If you promise never to leak a stop string to the client, you cannot release any suffix of the buffer that is a proper prefix of some stop string, because the next token might complete it. So the release point trails the generation point by up to max_stop_length - 1 characters. That single constraint is the entire latency and buffering story of this stage. The classical algorithm for the underlying problem is Aho-Corasick — a trie over the pattern set with failure links, giving amortised constant work per character and the longest match ending at each position — though most engines just scan a bounded tail, because a stop set of four short strings does not justify the machinery.
The token-boundary problem is the part that actually bites. Tokenizers are ambiguous encoders: a double newline may be one token in one vocabulary and two in another, and sequences like a quote followed by a comma and a newline are frequently a single merged token. All of those decode to text containing your stop string, which is why matching on decoded text is correct and matching on token ids is not. stop_token_ids is a separate, deliberately narrower mechanism: exact, cheap, and blind to any other encoding of the same characters.
Finally, stop sequences sit alongside the model's learned end-of-turn token, not in place of it. A chat model already knows how to stop. Stop strings are for structure inside a turn — few-shot delimiters, ReAct-style Observation markers, section boundaries in a completion-style prompt.
At a glanceSee it
Generated tokens are decoded, scanned for stop strings, and held back until a match is ruled out.
The knobsHyperparameters and nuance
- stop / stop_sequencesa list of strings. OpenAI's Chat Completions API caps it at four; Anthropic exposes
stop_sequencessimilarly. Each additional string, and each longer string, widens the holdback window and makes streaming choppier. - stop_token_ids(vLLM) — matched pre-detokenisation against the sampled id. Exact and free, but it fires only on that specific id, so it misses every other tokenisation of the same text. Use it for control tokens, not for prose delimiters.
- include_stop_str_in_output(vLLM) — default false. Set it true whenever you intend to re-feed the generated text into another prompt, otherwise the delimiter your format depends on has been silently removed.
- ignore_eos(vLLM) — keeps decoding past the end-of-sequence token. Essential for fixed-length throughput benchmarking, catastrophic in production, and a genuinely common copy-paste accident from benchmark scripts.
- min_tokens(vLLM) / min_new_tokens (HuggingFace) — the fix for a model that terminates immediately because the few-shot prefix already looks complete to it. Know the scope, because it is narrower than the name suggests: vLLM's
min_tokensmasks EOS andstop_token_idsin the logits until the floor is reached, and HuggingFace'smin_new_tokensblocks EOS. Neither suppresses stop strings, which are matched after detokenisation, so a stop string can still end the generation below your floor. - max_tokensthe other terminator, and the one you must distinguish in logs. A run where
finish_reasonislengthhalf the time is a different failure from one where it isstop.
EffectHow this stage moves the answer
The visible effects are truncation and lumpiness. A stop string that is too generic — a bare newline, a bare quote — cuts valid answers short, and because the string is stripped, the output reads as though the model simply gave up. Users interpret that as laziness. In few-shot and agent-style prompts the opposite failure is worse: with no stop configured, the model sails past its answer and hallucinates the next example, including a fabricated Observation block it never observed, and a downstream parser happily accepts that as real tool output. On the streaming side, a long stop string forces the server to withhold characters until termination, so a short answer can appear to hang and then arrive all at once. None of this is model behaviour. All of it is string matching, and all of it is configuration you own.
EvalsWhat it does to your measurements
Silent truncation is the eval failure that leaves no fingerprint in the score's shape. A stop of a single newline on a chain-of-thought benchmark cuts every answer at the first line break; accuracy collapses and the outputs look exactly like a model that will not reason. Always log finish_reason and report its distribution: a healthy generative eval is mostly natural EOS, with length rare and explicit stop-string hits at whatever rate your prompt format implies. Two harnesses evaluating the same model with different stop lists are not comparable and should not be put in the same table. And when a benchmark's reference implementation ships stop strings for its few-shot delimiters, dropping them changes both the score and the token bill, because the model keeps generating examples nobody reads. Check the stop configuration before you believe any generative benchmark number, including your own.
Failure modesWhen it goes wrong
- Answers end abruptly mid-sentencean overly generic stop string, most often a single newline inherited from a completion-style prompt template.
- Output is missing its final punctuationthe stop string shared a token with real content and the implementation dropped the entire token rather than splitting it.
- The stop never firesconfigured as
stop_token_idswhen the text was produced by a different token, or the string appears with different whitespace or casing than configured. - The model continues past the answer and invents the next few-shot exampleno stop configured for a completion-style prompt whose format implies one, so a parser downstream ingests fabricated tool output.
- Streaming stalls and then dumps everything at oncethe holdback window equals the longest stop string, so short answers are released only at termination.
PapersWhere this comes from
There is no research literature on this stage, and reaching for some would be dishonest — it is string matching over a stream, with a tokenizer-shaped complication. The normative sources are the provider API references, which define the semantics and the caps: OpenAI's stop parameter and Anthropic's stop_sequences. The reference implementations worth reading are vLLM's stop-checking path, which handles the buffer truncation and include_stop_str_in_output, and llama.cpp's partial-stop-string search, which is the clearest small example of the holdback invariant. The underlying algorithm, if you want the general form, is Aho & Corasick's “Efficient String Matching: An Aid to Bibliographic Search” (CACM, 1975) — a genuine paper about multi-pattern matching, and emphatically not a paper about language models.