Home › What Happens After You Hit Enter › Incremental detokenization
Pipeline stage · Operate

Incremental detokenization

Turning token ids back into readable text one at a time, when a single character can be split across several tokens.

In one line

Byte-level tokenizers can split one character across several tokens, so a correct streaming server sometimes emits nothing for a token and two characters for the next.

Why you'd careThe thing you have already noticed

You stream a response containing an emoji, or Hindi, or Japanese, and for a fraction of a second the client shows a black diamond with a question mark inside it, which then corrects itself. Or the concatenation of your streamed deltas has an extra space before a word compared with the same response fetched non-streamed. Both come from here. Modern tokenizers are byte-level: a token is a sequence of bytes, and one character of UTF-8 can occupy three or four bytes spread across two or three tokens. Decode each token independently and concatenate, and you hand the client incomplete byte sequences, which the decoder replaces with U+FFFD. Doing it correctly requires carrying state across tokens.

In and outWhat goes in, what comes out

InA stream of sampled token ids, the tokenizer's vocabulary and merge rules, and the decoder's carried state: a short window of previously emitted token ids plus any incomplete UTF-8 byte sequence left over from the last step.
ProcessDecode a suffix window of recent token ids to text, decode the same window minus the newest id, and release only the difference. Bytes that do not yet form a complete code point are retained for the next token. Special-token ids are dropped or rendered according to configuration.
OutA text delta — possibly empty, possibly carrying more characters than the newest token alone contributed — plus updated decoder state. The client concatenates deltas and must end up with exactly the non-streamed string.

The token ids are the ground truth and they survive; text is a derived view. What is lost is the token boundaries themselves. Once you hold only the string, you cannot in general recover which ids produced it, because re-tokenising text does not reliably reproduce the original encoding. That matters when a system feeds the model's own output back in on the next turn and gets a different id sequence than the model actually emitted — which breaks prefix-cache hits and, in agent loops, changes what the model sees it said.

ConceptThe idea underneath

No neural network content here either. The systems idea is stateful stream decoding under an alphabet mismatch.

Vocabularies are built by merging frequent byte pairs, so the units are byte sequences, not characters. GPT-2-style byte-level BPE maps all 256 byte values into the vocabulary through a reversible surrogate alphabet, so any byte string is encodable. SentencePiece with byte fallback achieves the same by emitting single-byte tokens for anything unseen. Either way, a four-byte emoji may be three tokens, and one of them may be a lone continuation byte that is not valid UTF-8 standing alone.

The naive loop — decode each id and emit it — fails for three separate reasons. Incomplete UTF-8, as above. SentencePiece's word-boundary marker, which becomes a leading space only in context, so a token decoded alone gives a different string than the same token decoded in sequence. And post-processing rules such as removing a space before punctuation, which cannot be decided from one token in isolation.

The correct algorithm, which is what vLLM and TGI implement, keeps a sliding window of the last few token ids and emits decode(ids[0..n]) minus decode(ids[0..n-1]), computed over that suffix window rather than the whole sequence for efficiency. vLLM has used a fixed lookback of a handful of tokens. The window must be long enough to cover the longest context-dependent post-processing rule in the tokenizer.

The invariant that all of this exists to satisfy is one line: the concatenation of every delta must equal decode(all_ids) exactly. If your streaming and non-streaming code paths do their post-processing differently, that invariant breaks and nothing will tell you.

At a glanceSee it

Incremental detokenization diagram

A sliding window of token ids is re-decoded each step so only genuinely new characters are emitted.

The knobsHyperparameters and nuance

  • skip_special_tokens(HuggingFace, vLLM; default true) — hides control ids from the output. Convenient for display, but end-of-turn markers and thinking delimiters are real vocabulary entries, and hiding them breaks any parser that segments the stream on them.
  • spaces_between_special_tokens(vLLM; default true) — inserts a space between rendered special tokens. A recurring source of off-by-one whitespace differences between streamed and non-streamed responses when special tokens are not skipped.
  • clean_up_tokenization_spaces(HuggingFace tokenizers) — legacy post-processing that strips spaces before punctuation and contractions. The asymmetry to watch is a version boundary, not a fast-versus-slow one: it defaulted to True for both fast and slow tokenizers through transformers v4.44, was deprecated in v4.45 with a FutureWarning, and moves to False thereafter. Individual models also pin it in their tokenizer_config.json. Two services on different transformers versions, or two checkpoints with different configs, will therefore disagree on the same token stream.
  • detokenizer lookback window(engine-internal, exposed as a constant rather than a request parameter in vLLM) — too small and context-dependent rules misfire at window edges; larger costs CPU per token per sequence, which is charged on the server's critical path at every concurrency level.
  • detokenize=false / --skip-tokenizer-init(vLLM) — return raw token ids and detokenise client-side. Removes per-token CPU work from the server hot path, which is a real throughput win at high concurrency, at the cost of moving tokenizer version-matching into your client.
  • add_special_tokens(encode side) — the mirror-image setting. Getting it wrong when you round-trip the model's own output adds a spurious BOS token, shifting every position and destroying the prefix-cache hit you were counting on.

EffectHow this stage moves the answer

In English with a Latin-heavy vocabulary, essentially nothing goes wrong, which is why this stage is invisible right up until it is not. What degrades is non-English text, emoji, mathematical symbols, and code containing accented identifiers or box-drawing characters. The visible symptom is a replacement character in the middle of a word, or two words running together because a leading space went missing. The more dangerous version is silent: if your client concatenates deltas and the server's post-processing differs between the streamed and non-streamed paths, the string your application stores is not the string the model produced. Feed that back in on the next turn and the prompt no longer matches the cached prefix byte for byte, so you lose the prompt cache and pay full prefill on every subsequent message — a latency and cost regression whose root cause is three layers away from where it shows up.

EvalsWhat it does to your measurements

This is a whole-benchmark-invalidating bug that only appears on multilingual and code suites, which is why it survives so long. Exact match on a multilingual QA benchmark collapses if U+FFFD appears anywhere in the answer, and the result reads as a model that cannot speak the language rather than a decoder that cannot spell it. Two specific traps. First, harnesses that score the accumulated streamed transcript rather than the final message bake any streaming-only detokenisation bug directly into the score, and the same run scored from the non-streamed response gives a different number. Second, a harness that re-tokenises generated text to count output tokens will disagree with the server's own count, so every cost-per-correct-answer figure derived from it is wrong. The diagnostic is straightforward: score the same run from token ids instead of text. If the ids are right and the text is wrong, the model is fine and the decoder is not.

Failure modesWhen it goes wrong

  • A replacement character appears mid-stream and corrects itself a token latera multi-byte character split across tokens and decoded per token instead of across a window.
  • Words run together, or gain a spurious leading spacethe SentencePiece word-boundary marker or clean_up_tokenization_spaces applied inconsistently between the streamed and non-streamed paths.
  • Concatenated stream deltas do not equal the non-streamed responsetwo separate post-processing implementations behind one API, a difference no test catches unless it explicitly compares the two.
  • Thinking or tool delimiters are missing from the streamskip_special_tokens is hiding structural control tokens that a downstream parser needs to segment on.
  • CPU saturates before the GPU at high concurrencyper-token detokenisation running on the server's request path; move it off the hot path or return raw ids.

PapersWhere this comes from

  • Neural Machine Translation of Rare Words with Subword UnitsSennrich, Haddow & Birch (arXiv:1508.07909, 2015; ACL 2016). Introduced BPE for neural sequence models, which is why the model's units are merged byte sequences rather than characters or words. The entire detokenisation problem is a consequence of that choice.
  • SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text ProcessingKudo & Richardson, 2018 (arXiv:1808.06226; EMNLP 2018 demo). Defined the lossless round-trip design and the word-boundary meta symbol ▁ (U+2581), which is the direct source of the leading-space behaviour described above. Do not credit it with byte fallback: that arrived later as a library feature, --byte_fallback, added in SentencePiece v0.1.9, which reserves 256 byte symbols and decomposes an unknown token into its UTF-8 bytes. The paper itself maps unseen characters to an ordinary unknown symbol. The distinction matters, because whether your tokenizer has byte fallback determines whether a partial multi-byte character is representable at all.
  • RFC 3629, UTF-8, a transformation format of ISO 10646Yergeau, 2003. The normative rules for multi-byte sequences and for handling ill-formed byte sequences, which is what produces the replacement character. Incremental detokenisation itself has no research literature at all — the reference material is vLLM's incremental detokenizer implementation and the Unicode standard's conformance requirements.
A living map of modern AI — kept current every morning