Home › What Happens After You Hit Enter › Tokenization (Text → Token IDs)
Pipeline stage · Operate

Tokenization (Text → Token IDs)

The final translation step: the assembled prompt string becomes a list of integers, and from here on the model has no concept of characters or words.

In one line

The model never sees letters or digits, only integer ids for chunks of bytes, which is why it miscounts characters, fumbles long arithmetic, and prices the same sentence differently per vendor.

Why you'd careThe thing you have already noticed

You asked how many r's are in strawberry and got the wrong number from a model that can write a compiler. You noticed a Japanese or Tamil prompt costing two to three times what its English translation costs. You watched arithmetic that works on four-digit numbers fall apart on nine-digit ones. All three are this stage. By the time the model exists, the prompt is a list of integers pointing at chunks of bytes, and there is no character layer underneath — the letters inside a token are not addressable from within the network. Tokenization is also the last thing that happens in intake, not the first, despite where it usually sits in diagrams: normalization, context assembly and template rendering all run before it.

In and outWhat goes in, what comes out

InThe fully rendered prompt string emitted by the chat template, already normalized, plus reserved control-token markers such as the end-of-turn id, and any multimodal placeholders standing in for image or audio patches. One flat string, no role structure left.
ProcessA regex pre-tokenizer splits the string into candidate pieces — runs of letters, digit groups, punctuation, leading spaces. Each piece is then encoded by replaying a learned merge table greedily, or by Viterbi-decoding a unigram language model, with raw bytes as the fallback for anything unseen.
OutAn int32 sequence of ids drawn from a vocabulary of roughly 32k to 256k entries, its length in tokens, and a byte-offset map from each id back into the original string so streamed output can be re-emitted as valid UTF-8.

The string itself is recoverable — modern byte-level tokenizers round-trip losslessly, whitespace and case included. What is lost is addressability. The model can attend to the id for strawberry but has no operation that indexes its third character; letter-level facts have to be learned distributionally, from training text that happened to spell things out. The offset map preserves the mapping for your code, not for the network.

ConceptThe idea underneath

Byte-pair encoding starts with a base vocabulary of the 256 byte values, then repeatedly finds the most frequent adjacent pair in a training corpus, merges it into a new symbol, and records the merge. Do that 128,000 times and you have a merge table. Encoding replays those merges in order on new text, so the algorithm is deterministic and the vocabulary is closed: any byte sequence encodes, nothing is out-of-vocabulary.

SentencePiece's unigram model works the opposite way. Start from a large candidate set of pieces, estimate a probability for each, prune the ones whose removal costs the least likelihood, and at encode time choose the segmentation that maximises sum of log p(piece) over the pieces — a Viterbi search rather than a greedy replay. That makes multiple valid segmentations of the same string possible, which is what subword regularization exploits during training.

An id is only useful as an index. The embedding table is a matrix E of shape V x d_model; token id 4711 means row 4711. At the far end of the stack the LM head projects back with a d_model x V matrix, often tied to E. With V = 128,000 and d_model = 4096 that is roughly 525 million parameters per matrix — the reason vocabulary size is an architectural decision, not a formatting detail. Bigger vocabularies mean fewer tokens per document, so cheaper requests and more real text per context window, paid for in parameters and in a wider, slower final softmax.

The regex that runs before the merges matters as much as the merges. Llama 2 and Mistral split every digit into its own token; the tiktoken-style pattern used by Llama 3 groups digits up to three at a time. That one line decides whether 1234567 arrives as seven place-aligned symbols or as an arbitrary grouping the model has to realign, which is most of why long-number arithmetic is fragile.

At a glanceSee it

Tokenization (Text → Token IDs) diagram

How a rendered prompt string becomes token ids, with byte fallback and the embedding lookup that follows.

The knobsHyperparameters and nuance

  • vocab_size32,000 for Llama 2, 128,000+ for Llama 3 and current tiktoken vocabularies. Larger shortens sequences and cuts cost; push it too far and rare pieces are seen so seldom in pretraining that their embeddings stay near-random, while embedding and LM-head memory grows linearly.
  • byte_fallback(SentencePiece; on for Llama-family tokenizers) — encodes unknown characters as individual byte tokens instead of a single UNK. Off, and rare scripts or emoji collapse to UNK and are unrecoverable. On, and an unusual string can cost one token per byte.
  • digit splitting in the pre-tokenizer regexmodel-specific and not runtime-configurable. One digit per token gives clean place alignment for arithmetic; 1-3 digit groups shorten sequences but misalign columns. You cannot change this without retraining, only choose a different model.
  • allowed_special / disallowed_special(tiktoken) — decides whether a literal <|endoftext|> in user text becomes the reserved control id or ordinary text. tiktoken raises on disallowed specials by default; permit them and untrusted input can forge a turn boundary.
  • add_special_tokens / add_bos_token(Hugging Face tokenizers) — whether BOS is prepended. Leave it true after the chat template already emitted BOS and you get a double-BOS prefix, which measurably degrades several Llama-family checkpoints.
  • model_max_length / truncation_sidewhere a too-long input is cut. truncation_side="right" quietly removes the user's actual question when the evidence block above it is long.

EffectHow this stage moves the answer

Vocabulary and pre-tokenizer decide which questions are hard. Character-level tasks — count the letters, reverse the word, does this string end in ly — degrade because the model is reasoning about objects whose interiors it cannot inspect; it answers from memorised spellings, which is why it gets common words right and rare ones wrong. Arithmetic degrades at the length where digit grouping stops aligning with place value, and the failure looks like confident off-by-a-power-of-ten answers rather than refusals. Fertility decides truncation: at the same max_tokens, a Hindi or Thai answer is cut roughly where an English one would still have a third to go, so the visible symptom is mid-sentence stops in one language and clean completions in another. Under-trained tokens produce the strangest class of all — one rare string that makes the model evade, repeat itself, or answer about something else entirely.

EvalsWhat it does to your measurements

Token counts are the denominator of half the metrics you report. Tokens per second, cost per request and context utilisation are not comparable across models with different tokenizers unless you normalise per character or per completed task; a model with 20 percent lower fertility looks 20 percent slower on tokens/sec while finishing sooner. GSM8K and other arithmetic sets move with digit handling, multilingual suites move with fertility, and character-level probes move with almost nothing else. The quiet invalidation is truncation: harnesses that clip at model_max_length drop the tail of long prompts without raising, so a long-context eval silently becomes a short-context one and the scores look merely disappointing rather than broken. Pin the tokenizer version alongside the weights — vocabularies do get patched, and a changed merge table changes every cached prefix and every token count you previously recorded.

Failure modesWhen it goes wrong

  • Confidently wrong letter countsthe model has no character-level view inside a token and is recalling spellings rather than inspecting them.
  • Arithmetic that works at four digits and fails at ninedigit grouping in the pre-tokenizer misaligns place value across token boundaries.
  • Answers truncated only in some languageshigher tokens-per-character fertility reaches the same max_tokens ceiling much earlier.
  • A prompt that should hit the prompt cache misses every timea tokenizer revision or a special-token setting changed, so the id prefix differs even though the text is identical.
  • Bizarre, evasive output triggered by one specific rare stringan under-trained token whose embedding never received meaningful gradient during pretraining.

PapersWhere this comes from

  • Neural Machine Translation of Rare Words with Subword UnitsSennrich et al., 2016 (arXiv:1508.07909). Introduced byte-pair encoding as a text tokenizer, establishing the open-vocabulary scheme nearly every current LLM still uses.
  • Subword Regularization: Improving Neural Network Translation Models with Multiple Subword CandidatesTaku Kudo, 2018 (arXiv:1804.10959). Introduced the unigram language-model tokenizer and showed that sampling among alternative segmentations improves robustness; this is the algorithm behind SentencePiece's default mode.
  • Language Model Tokenizers Introduce Unfairness Between LanguagesPetrov et al., 2023. Measured tokens-per-text across many languages for the same tokenizers and showed the disparity translates directly into cost, latency and effective context length.
  • Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsLand and Bartolo, 2024. Gave a systematic method for finding vocabulary entries that exist in the tokenizer but barely in the training data, and documented the anomalous generation they cause.
A living map of modern AI — kept current every morning