Home › What Happens After You Hit Enter › Token embedding lookup
Pipeline stage · Operate

Token embedding lookup

Each token ID is swapped for a learned vector — a row lookup in a giant table, not a computation.

In one line

Nothing is computed at this stage; a row is copied out of a table, which is why embeddings cost gigabytes of memory and essentially no arithmetic.

Why you'd careThe thing you have already noticed

You added two special tokens to a tokenizer, forgot to resize the model, and got either a CUDA device-side assert or fluent gibberish. Or you compared two models both labelled 3B and found one carrying an extra gigabyte for no visible reason. Both land on the embedding table. It is a plain [vocab_size, d_model] matrix and a token ID is a row index into it — at a 256,000-token vocabulary and d_model 4096 in bf16 that is 2 GB of weights before layer 0 runs. No matrix multiply happens. A row is copied. That is why this stage never appears in a FLOPs budget and always appears in a memory one.

In and outWhat goes in, what comes out

InAn integer array of token IDs, length S, every value in the range 0 to vocab_size minus 1. During prefill S is the whole uncached prompt; during decode it is one ID per running sequence.
ProcessA gather. For each ID, copy row W_e[id] out of the [vocab_size, d_model] embedding matrix. Some architectures — the original Transformer, Gemma — then multiply the result by sqrt(d_model) so the embedding scale matches what the first normalisation layer expects.
OutA dense bf16 tensor of shape [S, d_model], the layer-0 residual stream. For a 4096-wide model and a 2,000-token prompt that is about 16 MB of activation, handed straight to the first attention block.

Order survives only as array index; there is no positional information inside the vectors yet, and RoPE has not run. What is lost is the surface string. Two byte sequences that tokenised identically are now the same vector and can never be told apart again. Detokenization at the far end reverses token IDs, not embeddings, which is why an ID with no sensible text still prints something.

ConceptThe idea underneath

A lookup table is a matrix multiply you refused to pay for. Formally the operation is one_hot(id) @ W_e: a vector of zeros with a single 1, multiplied by the [vocab_size, d_model] matrix. Every product is zero except the one row, so implementations skip the arithmetic and index directly. That equivalence matters because it explains where the numbers came from — W_e is trained by gradient descent like any other weight, so each row is whatever vector made the next-token loss smallest across the contexts that token appeared in.

The consequence is that the row is the model's entire prior about a token before any context arrives. A token seen ten million times in training has a well-shaped row. A token seen four times, or zero times because it survived tokenizer training but not corpus filtering, has a row close to its random initialisation. Those are the notorious glitch tokens: feed one in and the model emits something unrelated, because the vector handed to layer 0 points nowhere meaningful.

Most small models tie the input and output matrices: the final projection to logits is h @ W_e^T, reusing the same weights. That halves the vocabulary cost and forces a token's input meaning and its output prediction to share one representation. Larger models often untie them, spending a second [vocab_size, d_model] matrix so the two roles can diverge. On Llama 3.2 1B — hidden size 2048, vocabulary 128,256 — the embedding table is about 263 million parameters, roughly a fifth of the whole model, and untying would add the same again.

At a glanceSee it

Token embedding lookup diagram

Token IDs index rows of the embedding table to produce the layer-0 residual stream.

The knobsHyperparameters and nuance

  • vocab_sizefixed by the tokenizer, from 32k (Llama 2) to 256k (Gemma). Larger packs more characters per token, so documents cost fewer positions; it also enlarges the table, enlarges the logits tensor produced at every decode step, and adds undertrained rows. Too small and text fragments into many tokens.
  • hidden_sizethe width of every row, typically 2048 to 8192. Multiplies the table size linearly and sets the residual-stream width for the rest of the model.
  • tie_word_embeddingsHugging Face config boolean. True saves a full vocabulary-sized matrix and is standard below about 7B; false is standard above it, though Gemma ties at every size. Flipping it on a checkpoint trained the other way produces immediate garbage.
  • resize_token_embeddingsmust be called after adding special tokens. Skipping it is the single most common cause of an index-out-of-range crash on the first forward pass.
  • modules_to_not_convertGPTQ and AWQ configs conventionally exclude embed_tokens and lm_head from quantization; lm_head is in the default exclusion list. Including them saves real memory and reliably damages rare-token and non-English behaviour first.
  • tensor_parallel_sizevocabulary-parallel embedding shards the table by rows across GPUs and adds an all-reduce. The vocabulary is padded so it divides evenly across ranks, which is why Megatron exposes --make-vocab-size-divisible-by (default 128) and vLLM pads each rank's slice. That padding is a serving-side detail and is not where published vocabulary numbers come from: Llama 3's 128,256 is 128,000 BPE tokens plus 256 special and reserved tokens.

EffectHow this stage moves the answer

Vocabulary size is the lever here and it pushes answers in two opposite directions. A larger vocabulary packs more characters into each token, so the same document costs fewer positions — long inputs fit and the model's effective reach grows without touching attention. But every token added is a row that has to be trained, and the tail of a 256k vocabulary is thin. You see this as ragged behaviour on rare strings: unusual proper nouns, non-Latin scripts, base64 blobs, long digit runs. The model treats them as single opaque units it barely knows rather than as characters it can reason over, which is exactly why character-level tasks — counting letters, reversing a word, spelling something out — fail on tokens common enough to have been merged but rare enough to be undertrained. Nothing downstream repairs it; the vector was already wrong at layer 0.

EvalsWhat it does to your measurements

Two measurement traps live here. The first is perplexity: it is per-token, and tokens are not comparable across tokenizers, so a model with a larger vocabulary looks better at equal quality. Report bits-per-byte when comparing models whose vocabularies differ. The second is length and cost accounting. Any figure quoted in tokens — context consumed, cost per document, throughput — is measuring the tokenizer as much as the model, and the same corpus can differ 20 to 30 percent in token count between a 32k and a 256k vocabulary. Embedding itself almost never moves a quality score. If a benchmark regression appears after a vocabulary or checkpoint change, look for a skipped resize, a tokenizer and checkpoint version mismatch, or a quantizer that flattened the table — not for a subtle representational effect.

Failure modesWhen it goes wrong

  • CUDA device-side assert or index-out-of-range on the first forward passa token ID at or above vocab_size, almost always because special tokens were added to the tokenizer without resizing the model's embedding matrix.
  • Fluent output with the wrong content from the very first tokentokenizer and checkpoint from different versions; the IDs point at real rows, just not the intended ones.
  • One specific rare string reliably derails the modelan undertrained embedding row whose vector never moved far from its random initialisation.
  • Memory use well above the parameter count you expected on a small modeluntied embeddings plus a large vocabulary put both a table and a head in memory, and quantizers usually skip both.
  • Quality drops only on non-English or code inputs after quantizing4-bit embedding and head weights damage the rows with the least training signal first.

PapersWhere this comes from

  • Attention Is All You NeedVaswani et al., 2017. arXiv:1706.03762. Defined the learned input embedding shared with the output projection, and the sqrt(d_model) scaling that several current models still apply at this exact point.
  • Using the Output Embedding to Improve Language ModelsPress & Wolf, 2016. arXiv:1608.05859. Showed that tying the input embedding to the output projection improves quality while deleting a full vocabulary-sized matrix, which is why tie_word_embeddings exists and defaults to true on small models.
  • Megatron-LM: Training Multi-Billion Parameter Language Models Using Model ParallelismShoeybi et al., 2019. arXiv:1909.08053. Introduced vocabulary-parallel embeddings, the row-sharded layout every tensor-parallel server still uses and the reason vocabulary sizes are padded to round multiples.
A living map of modern AI — kept current every morning