How models work

Context Window

The maximum amount of text, measured in tokens, that a model can take into account at one time: instructions, conversation, documents and its own reply together.

THE CONTEXT WINDOW · example: 200,000 tokensSystem promptToolsConversation historyDocumentsRoom for the replyinput: re-read and billed on every requestoutputWhen the window fills, the oldest content drops out or is summarised; the model sees nothing outside it.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/context-window

In plain terms

Think of a desk of fixed size. Everything the model works with has to lie on it at once: your instructions, the conversation so far, any documents, and the page it is writing. When the desk is full, something has to come off before anything new goes on. What is not on the desk does not exist for the model.

Why it matters

The window sets what a model can do in one go: read a whole contract or only a chapter, hold a long working session or lose the thread. Bigger windows help, with three caveats: every token in the window is paid for on each request, responses get slower, and accuracy tends to drop as the window fills.

Example

A 300-page tender document runs to roughly 150,000 tokens. With a 200,000-token window it fits whole, leaving room for questions and answers. With a 32,000-token window it has to be split, searched or summarised first, and the design of the application changes accordingly.

Most often confused with

Context Window vs. Memory

Context WindowWhat the model can see right now
MemoryWhat is stored outside and brought back when needed

The context window is working space that empties when the conversation ends. Memory is a separate store, a file or a database, that persists between sessions. A model “remembers” something only when the application copies it from that store back into the window.

Under the hood

The limit covers input and output together, and providers usually cap output separately. Models are stateless: on every turn the application resends the entire conversation, so long chats are re-billed each time unless prompt caching applies. Window sizes range from a few thousand tokens in small models to hundreds of thousands or a million in large ones. Use degrades before the limit: models attend less reliably to material in the middle of long contexts (“lost in the middle”) and overall quality declines as contexts grow (context rot). Techniques for staying within budget: retrieval in place of bulk loading, compaction, subagents with their own windows, and external memory.

Written by Mehmet Erkek · Last updated: