Home › What Happens After You Hit Enter › Rotary position embeddings (RoPE)
Pipeline stage · Operate

Rotary position embeddings (RoPE)

Position is injected by rotating each query and key by an angle proportional to its index, so scores depend only on relative distance.

In one line

RoPE stores position as an angle rather than a vector, so scores depend only on the distance between two tokens, and extending context means rescaling angles rather than retraining.

Why you'd careThe thing you have already noticed

You pointed a serving stack at a model, set the context window to the number on the model card, and watched answers stay perfectly fluent while quietly becoming wrong past a certain length. Or you ran a 128k-context model that aced a lookup at 4k and missed the same fact at 60k. That boundary is RoPE. Position is not stored in the residual stream; it is applied as a rotation to every query and key at an angle proportional to the token's index. Angles the model never saw in training put queries and keys in configurations it has no learned response to. Nothing errors. Grammar is a local phenomenon and survives. Retrieval across distance is what breaks.

In and outWhat goes in, what comes out

InQuery and key tensors of shape [batch, heads, S, d_head], plus each token's absolute position index. In decode that index is simply the current cache length, which is why an off-by-one in cache accounting misplaces every subsequent token.
ProcessSplit each head vector into d_head/2 coordinate pairs. Rotate pair i by angle m * theta_i, where m is the position and theta_i = base^(-2i/d_head). Low-index pairs rotate fast, high-index pairs take tens of thousands of tokens to complete one turn.
OutRotated Q and K with identical shapes and identical norms. V is untouched — position never enters the value path, only the score that decides how values get mixed.

Rotation is orthogonal, so every vector's length survives exactly and only direction changes. Absolute position becomes unrecoverable from the score, which is the point: the dot product depends on m - n. The practical cost is that a key written into the KV cache has its position baked in. A cached prefix cannot be reused at a different offset without recomputation, which is why prompt caches only ever match from the very start of a prompt.

ConceptThe idea underneath

Earlier transformers added a position vector to the embedding. RoPE multiplies instead. Take a query vector for one head, read it as d_head/2 two-dimensional points, rotate each point by an angle that scales with the token's position, do the same to every key, then compute attention as usual.

The trick is what rotation does to a dot product. Rotating two vectors by the same amount leaves their dot product unchanged, so rotating a query at position m and a key at position n leaves a score that depends on m - n and nothing else: dot(R_m q, R_n k) = g(q, k, m - n). Absolute position goes in, relative position comes out. The model learns distance, not location.

The frequencies come from a single number. With theta_i = base^(-2i/d_head) and the conventional base of 10,000, the first coordinate pair completes a turn every few tokens and the last takes tens of thousands. The slow pairs are what encode long-range order, and they are exactly the ones that never completed a cycle inside the training length. Push past that length and those pairs enter angular territory with no training signal behind it.

Every context-extension method is a way of keeping the angles familiar. Position interpolation divides m by a scale factor so 32k positions compress into the 8k range the model knows. NTK-aware scaling and YaRN instead raise the base, stretching the slow wavelengths while leaving the fast ones near their trained behaviour, which preserves local resolution. Llama 3 skipped the retrofit entirely and trained with a base of 500,000.

At a glanceSee it

Rotary position embeddings (RoPE) diagram

Position enters as a rotation of Q and K, so attention scores depend only on token distance.

The knobsHyperparameters and nuance

  • rope_thetathe base frequency. 10,000 is the classic default; Llama 3 ships 500,000 and several long-context Qwen and Mistral variants use 1,000,000. Raising it makes distant positions more distinguishable and slightly blurs local ordering. Leaving it at 10,000 while advertising a long window gives fluent output with no long-range recall.
  • rope_scalingthe Hugging Face config object, for example rope_type = linear, factor = 4.0, or the yarn and dynamic variants. Static linear scaling costs measurable accuracy at short context because it applies to every request; dynamic variants only engage past the original length.
  • max_position_embeddings and original_max_position_embeddingsthe requested length versus the trained one. Their ratio is the scale factor. Raising the first without setting the second is the most common long-context misconfiguration there is.
  • partial_rotary_factorthe fraction of each head's dimensions that get rotated, 0.25 to 0.5 in GPT-NeoX and Phi families. The unrotated remainder carries position-free content; lower values leave more capacity for content and less positional precision.
  • rotary layout, interleaved versus split-halfGPT-J style pairs adjacent dimensions, GPT-NeoX style pairs dimension i with dimension i plus d/2. Both are called RoPE and they are not interchangeable. Mixing them produces fluent nonsense with no error anywhere.

EffectHow this stage moves the answer

Get this stage right and you gain reach; get it wrong and you get confident, well-formed answers that reference the wrong part of the document. The failure is silent because local coherence never depends on the slow-rotating coordinate pairs — grammar, style and paragraph flow all come from nearby tokens. What degrades is anything that requires locating one specific span among many: pulling a value out of a long config, citing the right section, tracking which of twelve files a function lived in. Raising the base trades a little local sharpness for long-range separability, which is why aggressively extended models sometimes get slightly worse at short-context editing even as their needle scores improve. A layout mismatch is different in kind: output is fluent from the first token and unrelated to the prompt, because every score was computed from a garbled position.

EvalsWhat it does to your measurements

RoPE is what long-context benchmarks are actually measuring. Needle-in-a-haystack tests and RULER move directly with theta and scaling settings, and their signature is diagnostic: accuracy holds flat and then falls off a cliff at a specific token count rather than degrading smoothly. If you see the cliff, go find the trained length. The trap is the mirror image — almost every standard benchmark runs under 4,000 tokens, so a broken or over-aggressive RoPE configuration passes MMLU, GSM8K and your regression suite untouched and only fails on the one customer who pastes a large file. Any evaluation of a context-extension change must place probes at multiple depths across the full claimed window and report accuracy per position rather than averaged, because averaging hides exactly the cliff you are looking for.

Failure modesWhen it goes wrong

  • Coherent output up to roughly the trained length, then loops or degenerationmax_position_embeddings raised without any rescaling, so positions past training produce angles the model has no response to.
  • Fluent text unrelated to the prompt from the first tokeninterleaved versus split-half rotary layout mismatch, usually introduced during a checkpoint conversion.
  • Reusing a cached prefix at a new offset gives wrong answerskeys are stored after rotation, so their position is baked into the cache and cannot be shifted afterwards.
  • Short-context quality drops after enabling long contexta static linear scaling factor applied to every request instead of a dynamic one that engages only past the original length.
  • Two servers give different long-context quality from the same weightsone is reading rope_scaling from config.json, the other from a converted checkpoint that dropped the field.

PapersWhere this comes from

  • RoFormer: Enhanced Transformer with Rotary Position EmbeddingSu et al., 2021. arXiv:2104.09864. Introduced rotary embeddings and proved the key property, that rotating Q and K by position leaves an inner product depending only on relative distance.
  • Extending Context Window of Large Language Models via Positional InterpolationChen et al., 2023. arXiv:2306.15595. Showed that compressing positions into the trained range, rather than extrapolating past it, extends context with minimal fine-tuning — the origin of the linear scaling factor.
  • YaRN: Efficient Context Window Extension of Large Language ModelsPeng et al., 2023. arXiv:2309.00071. Scaled the wavelengths unevenly, leaving high-frequency pairs alone, which recovers the local precision that uniform interpolation gives up.
  • Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationPress et al., 2021. arXiv:2108.12409. The main alternative, an additive distance penalty on attention scores that extrapolates by construction; useful as the contrast that explains what RoPE trades away.
A living map of modern AI — kept current every morning