Short-term memory is a decision about what survives the turn; the context window is only the budget that decision has to fit inside.
ConceptWhat it is
Short-term memory is the working state an agent keeps while a single task is in progress: what it has already tried, what came back, what it decided, and what it still intends to do. It is scoped to the task and it is expected to be thrown away when the task ends.
It is worth separating from the context window, because the two get used as synonyms and they are not. The window is a hard budget measured in tokens. Short-term memory is a policy about what deserves to occupy that budget. An agent with an enormous window and no policy fills it with its own transcript and degrades; an agent with a modest window and a good policy stays coherent for far longer, because every step it carries forward has earned its place.
How it worksThe mechanics
The state lives outside the prompt and is rendered into it each turn. That inversion is the whole mechanism: rather than appending to a conversation until it overflows, the agent holds a structured record — the goal, the plan, the results so far, the current step — and re-renders as much of it as the budget allows, most relevant first.
What that makes possible is eviction with a rule. Tool output is usually the largest and least reusable thing in the loop, so it is summarised to its finding and the raw payload dropped. Repeated failures collapse into one note saying the approach did not work, which is the only part a later step needs. The goal and the plan are pinned and never evicted, because an agent that loses those does not fail loudly — it carries on producing plausible steps toward a task it can no longer state.
At a glanceSee it
State is held outside the prompt and rendered into it each turn, so eviction becomes a rule rather than an overflow. The goal and plan are pinned; tool output is summarised to its finding.
When to use itWhere it fits
- In any agent that takes more than one step, which is what makes it an agent rather than a single call.
- When tool output is large relative to the window — the usual case, and the one where summarising pays immediately.
- On long-running tasks, where the loop will outlive any fixed transcript.
- Whenever an agent needs to avoid retrying something it has already tried and failed.
When NOT to use itLimits & anti-patterns
- For anything that must outlive the task; that is long-term memory and it carries a different set of risks.
- As a reason to use a bigger window instead of a policy — a bigger budget spent badly buys a longer decline.
- Summarising the goal or the plan, which is the one thing that must survive verbatim.
- In a single-call system, where the prompt IS the state and this is machinery with nothing to hold.
Trade-offsAdvantages & costs
Advantages
- Keeps an agent coherent well past the point where an appended transcript would overflow.
- Makes eviction a stated rule rather than whatever the window happens to truncate.
- Stops repeated attempts at an approach that has already failed.
- Costs fewer tokens per step, because the prompt carries findings rather than payloads.
Trade-offs & costs
- Summarising loses detail, and the detail lost is sometimes the one a later step needed.
- The eviction policy is real logic that has to be designed, tested and maintained.
- A structured state is harder to inspect at a glance than a plain transcript.
- Pinning too much reintroduces the overflow it exists to prevent.
ExampleIn the real world
A research agent runs twelve steps. Left as an appended transcript it carries every raw tool payload forward and by step eight the goal has scrolled far enough up that the model begins optimising for the most recent sub-question instead of the task. With the goal pinned and each tool result reduced to its finding, the same run finishes inside a fraction of the budget and the twelfth step is still visibly working on the thing that was asked in the first one.
ToolsHow to implement it
- A structured state objectgoal, plan, results, current step, rendered per turn rather than appended.
- A summariser on tool outputthe single highest-value eviction in most loops.
- Pinned fieldsthe goal and plan, never evicted at any budget.
- A step ledgerwhat has been tried and what it returned, so failure is not repeated.
Cost & effortWhat it takes
This usually reduces cost rather than adding it: carrying findings instead of payloads is fewer input tokens on every subsequent step, and the saving compounds with loop length. The spend it adds is the summarising call itself, which is small and bounded. The effort is in the eviction policy, and that is design work rather than infrastructure.
What changedWhat changed here
Updated this page A system that spawns fresh Claude Code instances pre-loaded with prior-session knowledge demonstrates one way to give coding agents long-term memory.
Three kinds of claim, strongest first. Signal runs every morning.