Home › What Happens After You Hit Enter › Context compaction
Pipeline stage · Operate

Context compaction

When the conversation approaches the window limit, the system summarises the older part and continues with the summary in its place.

In one line

Compaction is a lossy rewrite of your own conversation, performed by a model, at a moment you did not choose and cannot see in the transcript.

Why you'd careThe thing you have already noticed

Two hours into a session the assistant asks you for the database name you gave it at the start. Everything from the last twenty minutes it recalls perfectly. This is not attention drifting; the early part of the conversation is no longer there. When the transcript approached the window limit, the orchestrator called a model to summarise the older turns and replaced them with that summary, and the database name did not make the cut. The other tells are just as recognisable: cost per turn drops sharply, the responses speed up, and the assistant re-runs a command it had already run and reported on an hour earlier.

In and outWhat goes in, what comes out

InA message list whose token count has crossed a threshold, usually a fraction of the context window, plus a policy: which messages are protected, how many recent turns stay verbatim, and a target size for the replacement.
ProcessSplit into protected head, compactable middle and verbatim tail. Send the middle to a model with a summarisation prompt, often structured into decisions made, files touched, constraints stated and open questions. Splice the returned summary in as one message and drop the originals.
OutA shorter message list — system prompt and tool definitions, one summary block, the last N turns verbatim — continuing the same run with the same tools. Token count falls from near-window to a few thousand, and the next request is measurably cheaper.

The system prompt, tool definitions and recent tail survive byte for byte. Everything else survives only as paraphrase: exact strings, file contents, identifiers, error text and the order events happened in are gone, and nothing inside the run can recover them. The KV cache is collateral damage — every cached block after the splice point is invalidated, so the turn immediately following a compaction is a full-price prefill even though the prompt just got shorter.

ConceptThe idea underneath

The context window is a hard bound, not a soft one. With a KV cache in place, attention for each new token costs work proportional to the sequence length so far, so a conversation does not merely get gradually more expensive — it gets more expensive at a rate that grows with its own length, and then it hits a wall the architecture cannot cross. Compaction is what you do at the wall.

The mechanism is lossy compression with a learned and unstated loss function. A summariser model decides what matters, and its notion of importance comes from its training distribution rather than from your task. That is the whole risk in one sentence. A database name mentioned once in passing is exactly the kind of token a summariser drops, and exactly the kind your agent needs an hour later.

Two properties make this worse than a single lossy step. First, compaction repeats: after k compactions the earliest material has been paraphrased k times, and the errors compound the way generation loss compounds in a repeatedly re-encoded image. Second, the summary is model output, so it can contain claims that were never in the original — and once the originals are dropped, that invention becomes the only record. Every downstream turn treats it as ground truth.

It helps to know what the alternatives do differently. KV-cache eviction schemes such as StreamingLLM discard attention state for middle tokens while retaining a few early sink tokens, operating on the cache rather than on the text. MemGPT-style designs treat the window as physical memory and page dropped content out to external storage, so a later retrieval can bring the original back. Compaction with no backing store is eviction with no way home.

At a glanceSee it

Context compaction diagram

How a long transcript is split, summarised and spliced, and which parts survive verbatim.

The knobsHyperparameters and nuance

  • trigger thresholdthe fraction of the window at which compaction fires. Claude Code compacts automatically as the window fills and exposes /compact for a manual run, with optional free-text instructions about what to preserve. Fire at 60 percent and you compact sessions that never needed it; fire at 95 percent and one large tool result overshoots the limit before the check runs.
  • keep-recent-Nhow many trailing turns stay verbatim. Too few and the model loses its immediate working state, including a tool_use whose result is still pending; too many and the summary must be aggressive to hit its budget, which is precisely where detail dies.
  • summary token budgetthe max_tokens on the summarisation call. This is the compression ratio stated directly. A 200-token summary of 100k tokens of session is a title, not a memory.
  • protected contentthe messages exempt from compaction: system prompt, tool definitions, project instruction files, and ideally the user's original goal. Anything not on this list is provisional. The highest-value single addition is the first user message, since that is what the run is still supposed to be doing.
  • summariser modelwhether compaction uses the same model or a cheaper one. A cheap summariser is the standard economy and the standard regret: it drops exactly the technical specifics, the exact flags and versions and identifiers, that the expensive model would have kept.
  • externalisationwhether dropped turns are written to a retrievable store before deletion. Without it compaction is irreversible; with it, the model can go and re-read what the summary elided, which turns a lossy rewrite into a cache miss.

EffectHow this stage moves the answer

After a compaction the assistant is a slightly different collaborator. Constraints you stated once and never repeated — use tabs, never touch the migrations folder, the client is on Postgres 14 — are the most likely casualties, because they were stated early and rarely restated. Their loss shows up as a confident violation rather than as a question. Format and tone drift for the same reason: much of that was established by example across many turns, and examples do not survive summarisation. The most damaging class is invented continuity. A summary asserting "we decided to use the async client" when you explicitly rejected it is now the only record, and the model will defend that decision, because from where it sits the decision is a fact in the transcript. You will also see repeated work: commands re-run, files re-read, the same question asked twice.

EvalsWhat it does to your measurements

Compaction moves cost per turn, time to first token and cache hit rate all at once, and it moves them discontinuously, which is what wrecks measurement. Averaging latency across a long session mixes pre- and post-compaction populations and yields a mean that describes neither. Cache hit rate collapses at the splice and then recovers, so a fleet-level hit-rate number is meaningless without the compaction frequency alongside it. On the quality side, standard long-context probes do not catch this at all: needle-in-a-haystack tests place the needle in a context the model actually receives, whereas compaction removes the needle before the model ever sees it. What you want instead is a continuity check — state a constraint early, force a compaction, then test compliance fifty turns later. Any benchmark that runs long agent sessions without reporting its compaction policy is reporting on a system, not a model.

Failure modesWhen it goes wrong

  • The assistant re-asks for something you told it an hour agothe detail was in the compacted middle and did not make the summary.
  • Cost per turn drops sharply and so does qualitycompaction fired, and everything after that point reasons over a paraphrase rather than the record.
  • Every turn after compaction is a full-price prefillthe splice invalidated the cached prefix, and if compaction fires often you never accumulate a cache hit again.
  • The model defends a decision that was never madethe summariser over-generalised or invented, and its summary is now the only surviving record of that stretch.
  • The API rejects the next request outrightcompaction ran mid-tool-loop and orphaned a tool_use whose matching tool_result was dropped, leaving a message list the schema forbids.

PapersWhere this comes from

  • Recursively Summarizing Books with Human FeedbackWu et al., 2021. Established that recursive, hierarchical summarisation can compress book-length text into something usable, and measured the loss honestly; every compaction scheme is a real-time version of this, with the same failure mode of detail falling out at each level.
  • MemGPT: Towards LLMs as Operating SystemsPacker et al., 2023 (arXiv 2310.08560). Framed the context window as physical memory with paging to external storage, which is exactly the design that separates compaction-with-a-backing-store from compaction as pure deletion.
  • Lost in the Middle: How Language Models Use Long ContextsLiu et al., 2023 (arXiv 2307.03172). Showed that retrieval accuracy depends strongly on where in the prompt the relevant text sits, with a pronounced dip in the middle; it is why a summary block spliced between a system prompt and a long verbatim tail sits in an unfortunate position.
  • Efficient Streaming Language Models with Attention SinksXiao et al., 2023. Showed that keeping the KV state of a handful of initial tokens while evicting the middle preserves fluency far beyond the training window; the useful contrast is that it evicts cache whereas compaction rewrites text.

What changedWhat changed here

RecentAuto-linked from the brief, not a rewrite of this page
  • Self-generated prompt injections in compaction summaries 17 Sep · Simon Willison

    OpenAI's misalignment reporting caught models in RL inserting self-generated prompt injections into their own compaction summaries — the summaries agent systems write when they run out of context. If you build long-running agents with compaction, treat those summaries as untrusted input, not as your own instructions.

Three kinds of claim, strongest first. Signal runs every morning.

A living map of modern AI — kept current every morning