Home › What Happens After You Hit Enter › Sliding-window / local attention
Pipeline stage · Operate

Sliding-window / local attention

Each token attends only to the last N tokens, so cost and cache stop growing with conversation length.

In one line

A window bounds the cache, not the reach: information still travels far by hopping layer to layer, but it arrives as a summary rather than a quote.

Why you'd careThe thing you have already noticed

You gave a model a long document, asked for a summary, and got a good one. You asked for the exact figure in the third table and got a confident, wrong number. The document fit, the context window accepted it, and the answer was still built from something other than the text. Sliding-window attention is one reason. In these models most layers attend only to the last W tokens — 4,096 in Mistral 7B, smaller still in recent Gemma releases — and anything older is not merely deprioritised, it has been dropped from those layers' caches entirely. What reaches the answer from further back is whatever survived compression through the layers in between, which preserves gist reliably and exact tokens unreliably.

In and outWhat goes in, what comes out

InQ, K and V for the current tokens, a window size W, and each layer's designation as local or global. Positions are absolute, and the mask is derived from them rather than stored.
ProcessFor local layers, mask out every key older than m - W; global layers attend to everything. The local layers' cache then behaves as a ring buffer of W positions per layer instead of a list that keeps growing.
OutAttention over at most W keys per local layer, and a per-layer cache whose size stops growing once the conversation passes W. The memory profile flattens instead of climbing linearly.

Locality is preserved perfectly: everything inside W is attended to exactly as under full attention. What is irreversibly lost is the token-level content beyond W in local layers. Once those keys are evicted there is nothing to recompute from short of re-prefilling the entire prompt, and the only remaining trace of the old text is whatever the residual stream is still carrying forward.

ConceptThe idea underneath

The idea is a locality prior borrowed from convolutional networks: most of what a token needs is nearby, so pay for nearby and let depth handle the rest. A single local layer sees W tokens. Stack two and the second layer's window covers tokens that already summarised their own windows, so the effective receptive field is 2W. With L local layers it reaches L * W — Mistral 7B's 32 layers at W = 4096 have a theoretical reach past 130,000 tokens.

Theoretical is the operative word, because reach is not recall. Information arriving via depth has been through L rounds of weighted averaging, and averaging destroys exactly the properties that matter for lookup: which token, in what order, spelled how. Gist survives compression; identifiers do not. It is the difference between remembering that a config set a timeout and remembering that it set it to 4500.

Production designs therefore mix. Longformer paired a local window with a small set of global tokens attending everywhere. Gemma 2 alternates local and global layers, so full-context lookups remain available at regular depths. Mistral 7B used windowed attention throughout with a rolling buffer cache. The ratio of local to global layers is the direct dial between memory and exact recall.

One more effect matters in practice. Models place large attention weight on the first few tokens of a sequence, an attention sink that absorbs probability mass when a query has no genuine match. Slide the window past those tokens and the softmax has nowhere to put that mass, and quality degrades sharply the moment the buffer starts rolling. Pinning a handful of initial tokens permanently fixes it.

At a glanceSee it

Sliding-window / local attention diagram

Local layers see only the last W tokens; global layers and depth carry information further.

The knobsHyperparameters and nuance

  • sliding_windowW, in tokens, exposed directly in Hugging Face configs and set to 4096 in Mistral 7B. Smaller W bounds the cache and speeds up long-context decode; it also sets the hard limit on exact recall in every local layer.
  • layer_types or the local-to-global interleaving ratiowhich layers see everything. One-to-one alternation is expensive and accurate; ratios of five-to-one or higher save far more memory and put much more weight on the few global layers doing the retrieval.
  • use_sliding_windowa boolean in Qwen 2 style configs that disables the mechanism entirely. Off means full attention, full cache and full recall, at full memory cost.
  • retained sink tokenshow many leading positions stay pinned in the cache regardless of the window. Zero produces a quality cliff the first time the buffer rolls; four is the usual value and is enough.
  • max_model_len versus the windowtwo separate limits, and the first is not evidence about the second. A server will happily accept 128,000 tokens into a model whose local layers see 1,024.

EffectHow this stage moves the answer

The degradation has a very specific shape: summaries stay good and quotes go bad. Ask what a long document is about and a windowed model performs like a full-attention one. Ask it to reproduce a value, a name, a version number or a line of code from early in that document and it produces something structurally right and factually invented — the correct kind of thing in the correct format with the wrong contents. Multi-turn conversation shows the same edge. An instruction given at the start, always answer in Spanish, never use bullet points, is followed for a while and then quietly abandoned once it slides out of the local layers, which reads as drift or boredom rather than memory loss. If a model obeys a constraint early and stops later without being told to, check the window before rewriting the prompt.

EvalsWhat it does to your measurements

This is the stage most likely to make a model look better than it is. Short-form benchmarks — MMLU, GSM8K, chat preference sets — run well inside W and are entirely insensitive to it, so a windowed model posts identical scores to a full-attention one while holding a fraction of its recall. What separates them is depth-resolved measurement: needle-in-a-haystack or RULER with the target placed at many positions across the full advertised window and scored per position rather than averaged. The signature is a plateau followed by a decline beginning near W. Two related traps: serving metrics look excellent because cache pressure is low and throughput is high, and any eval that puts its question immediately after the context is testing only the last W tokens no matter how long the context was. Vary the depth or you are not measuring the window at all.

Failure modesWhen it goes wrong

  • An instruction from the top of a long conversation stops being followedit has slid out of the local layers' window, and only its gist survives in the residual stream.
  • Quality falls off a cliff the moment context exceeds the windowno attention-sink tokens are pinned, so the softmax loses the positions that were absorbing unmatched probability mass.
  • The server accepts a 128,000-token request and answers as if it read a fraction of itthe advertised context is the position limit; the attention span is W.
  • Memory grows without bound despite sliding_window being setthe attention backend does not implement windowed masking and fell back to full attention, allocating a full cache.
  • Retrieval quality differs sharply between two servers running the same weightsone honours the config's window and interleaving, the other ignores both.

PapersWhere this comes from

  • Longformer: The Long-Document TransformerBeltagy et al., 2020. arXiv:2004.05150. Combined a sliding local window with task-specific global tokens, establishing the local-plus-global pattern production models still use.
  • Generating Long Sequences with Sparse TransformersChild et al., 2019. arXiv:1904.10509. Showed factorised sparse attention patterns retain quality at a fraction of the cost, the origin of structured attention sparsity.
  • Mistral 7BJiang et al., 2023. arXiv:2310.06825. Shipped sliding-window attention with a rolling buffer cache in a widely deployed open model, which is where the concrete W = 4096 numbers in this page come from.
  • Efficient Streaming Language Models with Attention SinksXiao et al., 2023. arXiv:2309.17453. Identified the attention-sink phenomenon and showed that retaining a few initial tokens prevents the quality collapse that otherwise hits when a window starts rolling.
A living map of modern AI — kept current every morning