The context window is how much text a model can hold in its head at once.
DefinitionWhat it means
The context window is the maximum number of tokens, spanning both input and output, that a model can process in a single call, ranging from a few thousand tokens in early models to over a million in current frontier systems. Anything beyond this limit must be truncated, summarized, or retrieved on demand, since the model has no memory of it within that request.
Why it mattersWhy you should care
Context window size determines what is actually possible in an application: a small window forces chunking and retrieval strategies like RAG, while a large window enables feeding an entire codebase or document set directly into a prompt. It also has real cost and latency implications, since processing cost scales with tokens used, so engineering teams weigh a bigger window against the benefit of simpler architecture.
At a glanceSee it
The window is one shared token budget — system prompt, chat history, retrieved docs and tool output all compete for the same space, including the slice reserved for the reply itself.
Long conversations survive by looping — when the window nears its limit the app summarizes old turns to reclaim budget, at the cost of details dropped from the compacted memory.
Where you see itIn the wild
- Model spec sheets listing a context window in tokens
- RAG systems built specifically to work around a limited window
- Long-document summarization use cases citing window size