Prompting & context

Context Rot

The decline in a model's accuracy and focus as its context window fills with more tokens, even when the limit has not been reached.

accuracytokens in context →short context: focusedlong context: details slipillustrative curveFitting in the window does not mean the model uses all of it with equal attention.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/context-rot

In plain terms

A person handed a five-page brief absorbs it well. Hand them five hundred pages and ask the same question, and they miss things, though every page was technically “read”. Models behave similarly. As the context grows, each piece of information gets a smaller share of attention, early instructions fade, and irrelevant material starts to interfere.

Why it matters

It overturns the comfortable assumption that a bigger context window solves the problem of giving a model enough information. More context can make results worse. For long-running agents it is the main cause of the familiar pattern: sharp for the first twenty minutes, then forgetful, repetitive and prone to ignoring instructions. Managing it is a design task and not something the model does for you.

Example

A research agent begins well. Forty tool calls later, its context holds 150,000 tokens of search results, most of them irrelevant. It starts re-running searches it has already done and forgets a constraint from the original brief. Restarting with a 2,000-token summary of findings and the original brief restores its performance at once.

Most often confused with

Context Rot vs. Context window limit

Context RotGradual loss of quality as the window fills
Context window limitThe hard maximum the window can hold

The limit is a wall: beyond it the request fails or content is cut off. Context rot is a slope that starts long before the wall. A model with a million-token window can show degraded recall and reasoning at a small fraction of that, depending on the task and on how much of the content is distracting.

Origin: The term gained currency in 2025, notably through a Chroma research report measuring the effect across many models.

Under the hood

Contributing effects: attention is spread over more tokens; position bias (“lost in the middle”), where material at the start and end is recalled better; distractors, meaning passages similar to the target but irrelevant, which do more damage than unrelated text; and accumulated tool output and failed attempts that the model continues to condition on. Degradation varies by model and is worse for tasks requiring reasoning across the context than for simple lookup. Countermeasures are the toolkit of context engineering: retrieve less and more precisely, clear or compact old tool results, delegate exploration to subagents, keep durable notes outside the window, and start fresh sessions with a summary.

Written by Mehmet Erkek · Last updated: