At a glanceContext Engineering
Designing what actually occupies the context window - budget, order, and what to leave out.
Where it sitsPrompting writes the instructions. RAG finds the documents. Neither decides what fits.
Context engineering is the discipline that owns the space between them: given a finite window and a request in flight, decide what goes in, what gets compressed, what gets dropped, and what that costs. It is the layer that actually breaks in production — and the one most teams never staff.
The name is newer than the problem. Prompt engineering assumed a short, hand-written string and a single turn. That assumption died the moment systems started retrieving documents, calling tools, and running for fifty turns. Anthropic's working definition is the useful one: context engineering is "the set of strategies for curating and maintaining the optimal set of tokens during LLM inference" — all the tokens, including everything that lands in the window without a human typing it. The prompt is now a small minority of the context, and it is not the part that fails.
The failure mode is specific and it is not "the model is dumb". It is that a system with a correct prompt, a correct retriever, and a 200K window returns a worse answer than the same system with a third of the input. That is not a bug in any one component. It is a budgeting failure, and the research on it is unambiguous.
The budgetThe window is a budget, not a container
Treat the context window like disk and you will fill it. Treat it like an attention budget and you will start making the right calls. The mechanical reason is the transformer itself: attention lets every token attend to every other token, which is n² pairwise relationships for n tokens. Double the input and you have quadrupled the relationships the model must resolve with the same trained capacity. Attention gets spread thinner. Nothing errors — the answer just gets worse.
There is a second, less-discussed reason: training distribution. Shorter sequences are far more common in training data than long ones, so models have proportionally less experience resolving long-range dependencies than the advertised window implies. The result, in Anthropic's framing, is "a performance gradient rather than a hard cliff" — precision decays smoothly, which is exactly why it escapes notice. A cliff shows up in your error rate. A gradient shows up as a system that is quietly mediocre.
The three rules that follow
- Every token competes.A token you add does not sit inertly beside the others — it dilutes attention across all of them. The question is never "does this fit" but "does this earn its place against everything already there".
- Advertised length is not usable length.The number on the model card is a structural limit, not a performance guarantee. The gap between the two is the single most expensive misunderstanding in applied AI (evidence below).
- Position is a variable you control.Where a fact sits changes whether the model uses it. This is free to exploit and free to get wrong.
Four findingsFour findings that should change how you build
Not vibes. Controlled studies, most of which isolate input length as the only variable.
1. Position effects are real, large, and survive bigger windows. Liu et al.'s Lost in the Middle (TACL 2024) placed the answer-bearing document at every position in a multi-document QA context and measured accuracy. Performance follows a U-shape: highest when the relevant information is at the very beginning or the very end, materially degraded when it sits in the middle. The result that should stop you cold is the comparison against a closed-book baseline. GPT-3.5-Turbo scored 56.1% with no documents at all and 88.3% with only the correct document (oracle). But in the 20- and 30-document settings, worst-case positioning drove accuracy below the 56.1% closed-book number. Retrieval made the system worse than not retrieving. The paper is explicit: "in the worst case, performance in 20- and 30-document settings is lower than performance without any input documents."
Two follow-on details matter for design. Extended-context models were not exempt — the degradation tracked the model, not the window size. And on a synthetic key-value retrieval task, model families diverged sharply: Claude-1.3 was near-perfect across 75, 140 and 300 key-value pairs, while GPT-3.5-Turbo and MPT-30B-Instruct showed pronounced U-curves. Position robustness is a per-model property you must measure, not a capability you can assume.
2. Reasoning degrades far below the advertised limit. Levy, Jacoby and Goldberg (ACL 2024) built FLenQA specifically to isolate input length: the same reasoning task, padded to roughly 250, 500, 1000, 2000 and 3000 tokens. Across GPT-4, GPT-3.5-Turbo, Gemini-Pro, Mistral 70B and Mixtral 8x7B, average accuracy fell from about 0.92 to 0.68 as input grew to 3000 tokens — degradation visible from around 500 tokens, and their conclusion is that models "quickly degrade in their reasoning capabilities, even on input length of 3000 tokens, which is much shorter than their technical maximum." Three thousand tokens. Against a million-token window. The paper also found that next-word-prediction quality correlates negatively with reasoning performance on their set — the standard metric points the wrong way.
3. Effective context is roughly a quarter of claimed context. Two benchmarks converge here. RULER evaluated 17 long-context models across 13 tasks and found that although models claim 32K or more, only about half maintain satisfactory performance at 32K — despite near-perfect scores on the basic needle-in-a-haystack test everyone quotes. NoLiMa went further and removed the crutch: it builds needles that share minimal lexical overlap with the question, so the model must infer a latent association rather than pattern-match a string. Of 13 models claiming at least 128K support, 11 dropped below 50% of their own short-context baseline at just 32K. Even GPT-4o, one of the best performers, fell from a 99.3% baseline to 69.7%. The mechanism the authors identify is precisely the one above: attention struggles to retrieve without a literal match to latch onto, and that struggle compounds with length.
This is the practical takeaway: needle-in-a-haystack is close to a marketing benchmark. It tests verbatim string retrieval — the easiest possible long-context task. Passing it says almost nothing about whether a model can reason over the same context. Do not accept it as evidence of long-context capability, and do not build your internal eval that way.
4. Degradation is non-uniform and partly counter-intuitive. Chroma's Context Rot report ran 18 models — Claude Opus 4 / Sonnet 4 / 3.7 / 3.5 / Haiku 3.5, o3, GPT-4.1 (plus mini and nano), GPT-4o, GPT-4 Turbo, GPT-3.5 Turbo, Gemini 2.5 Pro / Flash, Gemini 2.0 Flash, and Qwen3 at 235B / 32B / 8B — through controlled variations of input length. The headline: models do not use their context uniformly, and reliability decays with length even on trivial tasks. The specifics are more useful than the headline:
| Variable tested | Setup | Finding | What it means for your build |
|---|---|---|---|
| Needle-question similarity | 8 needle variants per haystack; similarity 0.445–0.775 (Paul Graham essays), 0.521–0.829 (arXiv papers) | Lower similarity degrades faster with length | Vocabulary-mismatched queries fail first and fail silently. Your hardest real queries are the low-similarity ones. |
| Distractors | Baseline vs 1 distractor vs 4 distractors | Even one distractor degrades performance; four degrade non-uniformly | Retrieval precision beats retrieval recall. A near-miss chunk is not free — it is actively harmful. |
| Haystack structure | Logically coherent prose vs randomly shuffled sentences | Models did better on the shuffled haystack | Counter-intuitive and instructive: coherent surrounding prose competes for attention. Structure is not automatically an asset. |
| Focused vs full input | LongMemEval, 306 prompts at ~113K tokens each, vs ~300-token focused versions | Large gap between focused and full; most pronounced on Claude models | The same question, same relevant facts, ~113K of surrounding context = measurably worse. This is the whole argument in one experiment. |
| Pure replication | 1,090 variants; repeat a word list with one unique word inserted; 12 lengths from 25 to 10,000 words | Performance consistently degrades with length; models under-generate or misplace the unique word | This task requires no reasoning at all. If copying degrades with length, nothing is immune. |
Two footnotes worth knowing before you cite this study: on LongMemEval, Claude Opus 4 refused 2.89% of tasks and GPT-4.1 refused 2.55%; GPT-3.5 Turbo was excluded from the repeated-words task after refusing 60.29% of it on content filtering. Refusal rates are a real confound in long-context evaluation and most write-ups omit them.
Four claimants on one budget
Every context window is a negotiation between four things that all believe they are essential. They are not equally essential, they do not decay at the same rate, and they have wildly different cost profiles.
| Claimant | Volatility | Cacheable | Typical failure | The discipline |
|---|---|---|---|---|
| System prompt | Frozen | Yes — fully | Grows by accretion: every incident adds a rule, nothing is ever deleted. Ends as 4,000 tokens of contradictory edge cases. | Aim for the right altitude — specific enough to guide behaviour, general enough to leave the model strong heuristics. Brittle if-then logic and vague platitudes fail at opposite ends of the same axis. Interpolating a timestamp or user ID here destroys the cache for everything downstream. |
| Tool definitions | Frozen | Yes — fully | Bloat. Twenty overlapping tools where the boundaries are ambiguous. | Anthropic's test is the honest one: if an engineer cannot say definitively which tool applies, the model cannot either. Tools must be self-contained, robust to error, and unambiguous about intended use. Adding a tool mid-session invalidates the cache from position zero. |
| Retrieved context | Per request | Rarely | Stuffing top-k because k is a config value nobody has revisited. Each marginal chunk is a distractor with a small chance of being useful. | Re-rank and cut hard. Lost in the Middle found retriever-reader performance plateaus well before retriever recall does — going from 20 to 50 retrieved documents bought roughly 1.5% for GPT-3.5-Turbo and 1% for Claude-1.3. You paid 2.5× the input tokens for one point. |
| History | Grows monotonically | Yes — as a prefix | The only claimant that grows without anyone deciding it should. Turn 40 carries every tool result from turn 3. | The only claimant that requires an active eviction policy. See compaction below. |
Note the tension the table exposes. The two cacheable claimants are the two that are cheap to keep, and the two volatile ones are expensive on every single call. The instinct to trim the system prompt for cost reasons is usually backwards: at cache-read rates it is close to free. The retrieved chunk you added "just in case" is not.
OrderingOrder is a design decision, and it is free
Given a U-shaped attention profile, position is a lever you are already pulling — the only question is whether you are pulling it deliberately. The rules fall out of the research directly:
Never bury the decisive fact in the middle. If your re-ranker produced an ordering, respect it — put rank 1 first or last, not in the centre of a ten-chunk block. If you have exactly one high-confidence chunk, consider sending only that: the oracle condition (88.3%) beat every multi-document condition in Liu et al.
Put the question last. The end of the window is the other high-attention position, and it is also the correct cache boundary — stable prefix first, volatile suffix last. Position and economics agree here, which is rare.
Restate the query around long context — but verify it helps. Liu et al. tested "query-aware contextualization" (placing the query both before and after the documents). On key-value retrieval it was transformative, taking models to near-perfect. On multi-document QA it "minimally affects performance trends". The same trick was a fix for one task and a no-op for another. This is the single best argument in this entire section for measuring on your own task instead of importing someone's tip list.
Cut distractors before you add context. Chroma found even a single distractor measurably degrades performance, and that failed responses disproportionately hallucinated from specific distractors. One near-miss chunk removed is worth more than one good chunk added.
Compaction, clearing, and notes
History is the claimant that grows on its own, so it is the one that needs a policy. There are three distinct mechanisms and they are routinely conflated. They are not interchangeable.
| Mechanism | What it does | Loses | Use when |
|---|---|---|---|
| Clearing | Prunes stale tool results outright, replacing each with a placeholder so the model knows content was removed. Anthropic's clear_tool_uses strategy defaults to a trigger at 100,000 input tokens, keeps the 3 most recent tool use/result pairs, and accepts exclude_tools plus a clear_at_least floor. | The raw results, permanently. Nothing is summarised. | Agentic loops where old tool output (file dumps, search results) is genuinely spent once read. Their token-count example shows 70,000 → 25,000 input tokens. |
| Compaction | Summarises the conversation near the window limit and reinitialises with the compressed version. Claude Code passes message history to the model to summarise, then continues with that summary plus the five most recently accessed files. | Detail, lossily and irreversibly. What survives is whatever the summariser judged important. | Long conversational or coding sessions that must continue past the window. Tune the summariser to preserve architectural decisions, unresolved bugs and implementation details while discarding redundant tool output. |
| Structured notes | The agent writes durable memory to an external file it can re-read later. | Nothing — the note outlives the window. | Anything that must survive a context reset. This is the only mechanism that is not lossy, because the state lives outside the budget entirely. |
The ordering that works: externalise first, clear second, summarise last. A note on disk costs nothing against the budget and loses nothing. Clearing is cheap and honest — it removes content the model has already metabolised. Summarising is the lossiest option and should be the last resort, because a summariser is an additional model call with its own failure modes, and every compaction cycle is a generation loss. Compact a conversation five times and you are reasoning over a summary of a summary of a summary.
Sub-agents are the fourth option and belong in the same family: a sub-agent explores with its own clean window and returns a condensed 1,000–2,000 token summary instead of its full trace. That is compaction with the compression pushed into a separate context. It works — but see the caveat in section 08 before you reach for it.
Cost pressureThe budget is also a bill — and the two constraints fight
Context engineering is usually taught as a quality problem. It is equally a cost problem, and the two have a conflict at the centre that nobody warns you about.
Prompt caching is a prefix match: providers hash the rendered prompt up to a breakpoint, and any byte change anywhere in that prefix invalidates everything after it. The economics are steep enough to dominate design. Cache reads cost roughly 0.1× base input price. Cache writes cost 1.25× for a 5-minute TTL and 2× for a 1-hour TTL. Break-even on the 5-minute cache arrives at just two requests (1.25 + 0.1 = 1.35 vs 2.0 uncached); the 1-hour TTL needs three, but survives gaps in bursty traffic.
The conflict, stated plainly
- Caching rewards an immutable prefix.Never touch the front of the window and you read at a tenth of list price forever.
- Context management mutates the prefix by definition.Clearing a tool result from turn 3 changes the bytes at turn 3, which invalidates the cache for turns 4 through 40.
- So every eviction has a cache-write cost attached.This is why Anthropic ships a
clear_at_leastparameter — it exists specifically to ensure you clear enough tokens to justify the invalidation you just caused. Clearing 500 tokens from a 90,000-token cached prefix is strictly worse than doing nothing. - Therefore: evict in large, infrequent batches, not continuously.The intuition to trim aggressively and often is the expensive one.
The corollary for architecture: put stable content first and volatile content last, always. Frozen system prompt, deterministically-ordered tools, then retrieved context, then history, then the question. A timestamp in the system prompt — datetime.now() in an f-string, a per-request UUID, an unsorted json.dumps — silently costs you the entire cache on every call, and it does so without an error, a warning, or a log line. The way you find it is by checking whether cache-read tokens are non-zero across repeated calls. If nobody is watching that number, nobody knows.
Bigger is not betterWhen a bigger window is the wrong answer
A larger window is a real capability and occasionally the right call. It is also the default reflex, and the default reflex is usually wrong. Here is how to tell.
| Symptom | The reflex | Why it fails | Do this instead |
|---|---|---|---|
| Model misses a fact that was in context | Add more context | The fact was already there. More context lowers the odds it is used — this is the lost-in-the-middle result exactly. | Move the fact to the front or the back. Cut the distractors around it. |
| Retrieval sometimes misses | Raise top-k | Plateau well before recall saturates: 20 → 50 documents bought ~1.5% / ~1%. Each added chunk is a distractor, and one distractor is enough to hurt. | Fix the retriever or add a re-ranker. Precision, not recall. |
| Long conversation degrades | Move to a 1M window | Degradation started around 3,000 tokens, not at the window limit. A bigger window postpones the hard failure and does nothing about the soft one you are actually experiencing. | Compact, clear, or externalise to notes. |
| "The model supports 128K, so 128K works" | Fill it | 11 of 13 models claiming 128K+ dropped below half their short-context baseline at 32K (NoLiMa). Only about half of 17 models held up at 32K on RULER. | Measure your effective context on your task. Assume roughly a quarter of the advertised number until proven otherwise. |
| One agent can't hold it all | Fan out to parallel sub-agents | Cognition's argument: parallel sub-agents each act on a partial view and make conflicting implicit decisions about style, edge cases and interpretation, with no shared global context to reconcile them. You traded a context problem for a coherence problem. | Default to a single-threaded linear agent with continuous context. If you must parallelise, keep writes single-threaded and let extra agents contribute intelligence rather than actions. |
The honest case for a bigger window is narrow and worth stating: a genuinely irreducible input that cannot be chunked without destroying meaning — a whole codebase for a cross-cutting refactor, a single long legal instrument where any clause may condition any other. There, the alternative is not "smaller context", it is "no answer". Take the degradation knowingly and build the eval to measure it. That is a different act from filling a window because it was there.
What to instrument
Context problems are invisible by construction. Nothing throws. The system returns a plausible, slightly worse answer, and the gradient is smooth enough that no alert fires. If you do not measure it, you will not see it — you will just have a product that quietly underperforms its own components.
The five numbers
- Effective context, measured on your task.Run your real eval at increasing input lengths and find where accuracy breaks. Not needle-in-a-haystack — a task that requires reasoning over the context and whose queries do not lexically match the source. That is the NoLiMa lesson: lexical overlap is a crutch that hides the failure.
- Position sensitivity.Take a passing case, move the decisive fact from the front to the middle, re-run. The delta is your U-curve. If it is large, ordering is a live bug in your pipeline and no amount of prompt tuning will touch it.
- Tokens by claimant.Break the window down: system, tools, retrieved, history, query. Most teams have never looked and are surprised by the answer — usually history or top-k retrieval eating the majority.
- Cache-read ratio.If cache reads are near zero across repeated calls with the same prefix, something volatile is sitting in the prefix. This is a pure cost leak and it is trivially detectable once someone looks.
- Compaction generation count.How many times has this session been summarised? Quality degrades per cycle. Nobody tracks it, and it explains a specific class of "the agent forgot what we decided" complaint.
One closing note on method. Almost every finding on this page is counter-intuitive in a specific direction: less context outperformed more. Shuffled haystacks beat coherent ones. One distractor was enough to hurt. Fifty documents barely beat twenty. Retrieval was sometimes worse than no retrieval. The engineering instinct — when the answer is wrong, give it more to work with — is consistently the wrong instinct here. Budget accordingly.
Sources
Every external claim on this page traces to one of these. Where a number is quoted, it is quoted from the paper, not from a summary of it.
Primary research and vendor documentation
- Liu et al. — Lost in the Middle: How Language Models Use Long Contexts (TACL 2024)
- Chroma Research — Context Rot: How Increasing Input Tokens Impacts LLM Performance (2025)
- Modarressi et al. — NoLiMa: Long-Context Evaluation Beyond Literal Matching (ICML 2025)
- Levy, Jacoby & Goldberg — Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of LLMs (ACL 2024)
- Hsieh et al. — RULER: What's the Real Context Size of Your Long-Context Language Models?
- Anthropic Engineering — Effective context engineering for AI agents
- Anthropic — Context editing (clear_tool_uses strategy)
- Anthropic — Prompt caching: pricing, TTL and cache economics
- Cognition — Don't Build Multi-Agents
Long-context evaluation moves quickly and model-specific numbers age fast — the position, dilution and effective-context effects have held across every generation measured so far, but re-run the benchmarks against the model you actually ship. Figures cited here reflect the models named in each study.
What changedWhat changed here
Nothing in the daily brief has touched this page since 2026-09-25. The sweep runs every morning and checks every page on this site; when it finds something for this one, it lands here.
Three kinds of claim, strongest first. Signal runs every morning.