In plain terms
A person handed a five-page brief absorbs it well. Hand them five hundred pages and ask the same question, and they miss things, though every page was technically “read”. Models behave similarly. As the context grows, each piece of information gets a smaller share of attention, early instructions fade, and irrelevant material starts to interfere.
Why it matters
It overturns the comfortable assumption that a bigger context window solves the problem of giving a model enough information. More context can make results worse. For long-running agents it is the main cause of the familiar pattern: sharp for the first twenty minutes, then forgetful, repetitive and prone to ignoring instructions. Managing it is a design task and not something the model does for you.
Example
A research agent begins well. Forty tool calls later, its context holds 150,000 tokens of search results, most of them irrelevant. It starts re-running searches it has already done and forgets a constraint from the original brief. Restarting with a 2,000-token summary of findings and the original brief restores its performance at once.
Most often confused with
Context Rot vs. Context window limit
The limit is a wall: beyond it the request fails or content is cut off. Context rot is a slope that starts long before the wall. A model with a million-token window can show degraded recall and reasoning at a small fraction of that, depending on the task and on how much of the content is distracting.
Origin: The term gained currency in 2025, notably through a Chroma research report measuring the effect across many models.
Under the hood
Contributing effects: attention is spread over more tokens; position bias (“lost in the middle”), where material at the start and end is recalled better; distractors, meaning passages similar to the target but irrelevant, which do more damage than unrelated text; and accumulated tool output and failed attempts that the model continues to condition on. Degradation varies by model and is worse for tasks requiring reasoning across the context than for simple lookup. Countermeasures are the toolkit of context engineering: retrieve less and more precisely, clear or compact old tool results, delegate exploration to subagents, keep durable notes outside the window, and start fresh sessions with a summary.