In plain terms
A model can only work with what is in front of it, and the space in front of it is limited. Context engineering is the editor's job: for this step of this task, what must the model see, in what order, and what should be kept out because it would only distract? A brilliant model with the wrong page open still gives the wrong answer.
Why it matters
As AI moved from single questions to agents that work for hours, the bottleneck shifted from phrasing to information supply. Most failures of capable models in production are context failures: the relevant fact was not retrieved, or it was buried under fifty irrelevant tool outputs. Teams that manage context well get more from a cheaper model than teams that do not get from the best.
Example
A coding agent keeps losing track in a large repository. The model stays the same. The fix is in what it sees: a short project guide loaded at the start, search results trimmed to the relevant functions, old tool output cleared after each subtask, and a running notes file it can reread. The same model then completes tasks it used to abandon.
Most often confused with
Context Engineering vs. Prompt Engineering
Prompt engineering asks “how should I word this?”. Context engineering asks “what should the model be looking at right now, and what should it not?”. The first is mostly done once, at design time. The second is a running system: retrieval, memory, compaction and tool design working throughout a task.
Origin: The term spread in mid-2025, popularised by Tobi Lütke and Andrej Karpathy among others, as a more accurate name than prompt engineering for building agent systems.
Under the hood
Levers: the system prompt and tool definitions (stable, cacheable); just-in-time retrieval in place of loading everything up front; tool results that return summaries and references in place of raw dumps; compaction of older turns; external memory and scratchpad files; subagents that work in separate windows and return digests; ordering, since position affects attention. The guiding constraint is that context is finite and quality falls as it fills (context rot), so the aim is the smallest set of high-signal tokens that lets the model do the next step.