Home › What Happens After You Hit Enter › Context Assembly & Ordering
Pipeline stage · Operate

Context Assembly & Ordering

System prompt, retrieved documents, conversation history, memory, and tool results are stitched into one sequence in a deliberate order — the last authoring decision before tokenization…

In one line

Position is a parameter that never appears in the API: the same fact placed at the top or the middle of a long prompt produces different answers.

Why you'd careThe thing you have already noticed

You added three more retrieved documents and the answers got worse. You confirmed a fact was in the prompt, watched the model claim it did not know, then moved the same sentence two thousand tokens earlier and got the right answer. You noticed a conversation that behaved well for ten turns and started drifting at twenty-five. None of these is a model problem. This is the stage where the system block, summarised history, verbatim recent turns, retrieved evidence, memory and the current question are stitched into one linear sequence, and the order is a design decision somebody made — often implicitly, often inside a loop that appends whatever it happens to have.

In and outWhat goes in, what comes out

InEvery context source with its token count and a volatility flag: the rendered system block, tool schemas, summarised earlier history, verbatim recent turns, the formatted evidence block, persistent memory, and the current user message.
ProcessChoose a section order, split the static prefix from the dynamic tail so the prefix can be cached, apply a total token budget with per-section caps, then drop or summarise the lowest-value sections until it fits. Tag each span's provenance and trust level.
OutOne ordered sequence, commonly two thousand to two hundred thousand tokens, carrying a marked prefix boundary for caching and a manifest recording which sources were included, truncated or dropped entirely.

Section boundaries are preserved only as text; nothing downstream enforces them. The drop decision is the irreversible part, and it is invisible to the model — no signal marks the place where three documents were removed, so the answer is produced with full confidence over a partial view. Your manifest is the only record; without it, a wrong answer is indistinguishable from a missing input.

ConceptThe idea underneath

Attention is softmax(QK^T / sqrt(d_k)) V. Read it as a weighted lookup: Q is the query vector for the token being computed, K holds a key for every earlier token, their dot products score relevance, sqrt(d_k) — the square root of the per-head dimension — keeps those scores in a range where softmax is not saturated, and V holds the values that get mixed. The consequence that matters here is the softmax itself: the weights sum to one. Attention is a fixed budget. Every token you add takes weight away from every token already present, so a prompt is not a container you can keep filling for free.

Empirically the budget is not spent evenly. Retrieval accuracy across a long context is U-shaped: material at the very start and the very end is used reliably, material in the middle much less, and the dip deepens as the context grows. Two mechanisms contribute. The first tokens act as attention sinks that many heads fall back on, and relative position encodings such as RoPE make attention between distant tokens harder to express.

Length degrades reasoning independently of retrieval. The same task, padded with irrelevant but harmless text, is solved less often; the model is not failing to find the input, it is failing to hold the whole problem. This is why the model has a 200k window and the model reasons well over 200k tokens are different claims, and only the first one is on the spec sheet.

Ordering also decides economics. Cache reuse is prefix-exact, so the sequence must be arranged stable-first: system block, tools, slowly-changing memory, then evidence, then the turn. An architecture that appends retrieved documents ahead of the system block passes every functional test and pays full prefill price forever.

At a glanceSee it

Context Assembly & Ordering diagram

Sections enter a token budget; over-budget assembly drops or summarises before emitting one ordered context.

The knobsHyperparameters and nuance

  • section ordera config decision with no standard parameter name. Two defensible defaults: stable-first for cache economics, and question-last so the instruction sits adjacent to the generation point. Putting the question first and the evidence last inverts the cache and buries the ask.
  • max_model_len(vLLM) or the provider's context limit — the hard ceiling. Set near the model's maximum and you trade KV cache capacity for sequence length, cutting how many requests run concurrently; set it low and long prompts either fail or get truncated.
  • --enable-prefix-caching(vLLM) and cache_control breakpoints (Anthropic) — these determine where the reusable prefix actually ends. Their placement must match the section order or neither one helps.
  • top_kfor retrieval — 3 to 10 is the usual productive band. Higher raises recall and dilutes attention; the crossover where added recall stops paying for added distraction is task-specific and worth measuring rather than guessing.
  • compaction thresholdthe fraction of the window at which history is summarised. Trigger too late and you truncate mid-task; too early and you discard detail the model still needed two turns later.

EffectHow this stage moves the answer

The characteristic degradation is not error, it is hedging. As the assembled context grows, answers get longer, more qualified and more summary-shaped: the model starts describing the material instead of using it, because a diffuse attention distribution produces a diffuse answer. Second, whichever section is longest sets the register — a prompt with eight thousand tokens of scraped documentation and two hundred tokens of instruction produces documentation-flavoured prose regardless of the instruction's tone. Third, burying the actual question under evidence produces answers to a nearby question, because the model latches onto the most salient recent topic, which is now the last document rather than the ask. Fourth, and worst, silent drops: the answer is confident, well-formed, correctly sourced from what remains, and gives no indication that the decisive paragraph was cut to make the budget.

EvalsWhat it does to your measurements

This stage usually dominates end-to-end quality in RAG and agent systems, often more than model choice does, and it is the stage least likely to be held fixed between runs. Long-context suites are the sensitive ones — needle-in-a-haystack probes for pure retrieval, RULER-style synthetic tasks for retrieval plus aggregation, and any multi-document QA set. Two silent invalidations. First, order nondeterminism: if assembly iterates a set or a dict whose ordering is not stable, every run shuffles the middle of the prompt and scores move by several points for no reason visible in the diff. Second, length mismatch: an eval built on two-thousand-token prompts says nothing about a production path that assembles sixty thousand, and the failure is not proportional, it appears abruptly past some length. Log the assembled prompt hash and its section manifest alongside every score.

Failure modesWhen it goes wrong

  • Adding more retrieved documents lowers accuracydistractors dilute a fixed attention budget faster than the extra recall pays for itself.
  • A fact verifiably present in the prompt is reported as unknownit landed in the middle of a long context, the weakest position.
  • Cost and latency rise while answers stay identicala volatile section sits ahead of the stable prefix, so nothing caches.
  • The model answers a question from two turns agothe current question is buried under a large evidence block and is no longer the most salient recent content.
  • Occasional confident answers that contradict a source you know was retrievedbudget-driven truncation dropped it, and nothing in the output marks the gap.

PapersWhere this comes from

  • Lost in the Middle: How Language Models Use Long ContextsLiu et al., 2023 (arXiv:2307.03172). Measured the U-shaped accuracy curve over document position and showed it persists in models explicitly built for long contexts, making position a first-class design variable rather than an artefact. The curve is a measured regularity; the paper offers no mechanism for it, and none is established in the literature.
  • Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language ModelsLevy et al., 2024 (arXiv:2402.14848). Held the task constant and varied only padding length, isolating context length itself as a cause of reasoning degradation well below the advertised window — measurable from roughly 3,000 tokens.
  • RULER: What's the Real Context Size of Your Long-Context Language Models?Hsieh et al., 2024 (arXiv:2404.06654). Built synthetic retrieval, multi-hop tracing and aggregation tasks at controlled lengths and found effective context far shorter than claimed context for most models; of 17 models all claiming 32K or more, only about half handled 32K effectively.
A living map of modern AI — kept current every morning