In plain terms
Think of a desk of fixed size. Everything the model works with has to lie on it at once: your instructions, the conversation so far, any documents, and the page it is writing. When the desk is full, something has to come off before anything new goes on. What is not on the desk does not exist for the model.
Why it matters
The window sets what a model can do in one go: read a whole contract or only a chapter, hold a long working session or lose the thread. Bigger windows help, with three caveats: every token in the window is paid for on each request, responses get slower, and accuracy tends to drop as the window fills.
Example
A 300-page tender document runs to roughly 150,000 tokens. With a 200,000-token window it fits whole, leaving room for questions and answers. With a 32,000-token window it has to be split, searched or summarised first, and the design of the application changes accordingly.
Most often confused with
Context Window vs. Memory
The context window is working space that empties when the conversation ends. Memory is a separate store, a file or a database, that persists between sessions. A model “remembers” something only when the application copies it from that store back into the window.
Under the hood
The limit covers input and output together, and providers usually cap output separately. Models are stateless: on every turn the application resends the entire conversation, so long chats are re-billed each time unless prompt caching applies. Window sizes range from a few thousand tokens in small models to hundreds of thousands or a million in large ones. Use degrades before the limit: models attend less reliably to material in the middle of long contexts (“lost in the middle”) and overall quality declines as contexts grow (context rot). Techniques for staying within budget: retrieval in place of bulk loading, compaction, subagents with their own windows, and external memory.