Prompting & context

Compaction

Replacing a long conversation history with a shorter summary so that an AI session can continue within its context window.

BEFORE · the window is nearly fullSystemLong conversation history and tool outputNewAFTER · compactedSystemSummaryNewfreed space: the work can go onThe summary decides what survives: decisions and open items stay, raw detail goes.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/compaction

In plain terms

A long working session fills the model's window: messages, files read, tool results, dead ends. Compaction is clearing the desk without losing the plot. The history so far is condensed into a summary that keeps what matters, such as decisions made, the current state and what remains to be done, and the session carries on from the summary with room to work again.

Why it matters

Compaction is what allows agents to work for hours on a task that would overflow any window. It is also a point where information is deliberately thrown away, so its quality matters: a summary that drops one constraint produces an agent that confidently violates it later. Anyone running long agent sessions should know when compaction happens and what it preserves.

Example

A coding agent has spent two hours refactoring a module and its context is 90 percent full. The system compacts: 180,000 tokens of history become a 4,000-token summary listing the files changed, the design decisions, the two failing tests and the next step. The agent continues, rereading from disk the three files it still needs.

Most often confused with

Compaction vs. Truncation

CompactionSummarises older content, keeping the meaning
TruncationDeletes older content outright

Truncation simply drops the oldest messages once the window is full: it is cheap and loses everything in them, including the original instructions if you are not careful. Compaction spends a model call to condense the history first. It costs more, and it keeps the thread.

Under the hood

Triggered at a token threshold or at natural breakpoints between subtasks. A model writes the summary under instructions on what to retain: goals and constraints, decisions with reasons, current state, open items, and references to files or records that can be re-fetched. Lighter variants remove only stale tool results and keep the dialogue (context editing). Trade-offs: summaries lose detail and can introduce errors, and replacing the history invalidates the prompt cache. Complementary techniques are external notes or memory files that survive compaction, and subagents that keep exploratory work out of the main context to begin with.

Written by Mehmet Erkek · Last updated: