In plain terms
A scientist running a long experiment keeps a lab notebook: what was tried, what happened, what comes next. After a weekend, or when a colleague takes over, the notebook is how the work continues without starting again. An agent needs the same, because its context window is its head, and in a long job that head is emptied or summarised many times. Agent memory is the notebook and the filing cabinet behind it: whatever the agent writes down so that it, or its successor, can pick up the thread.
Why it matters
Without memory, an agent's competence ends where its context ends: after a compaction or in a new session it repeats work, asks again and forgets what was decided. With it, an agent can carry a task across days and improve at recurring work by keeping what it learned. The cost is that memory has to be managed. A wrong lesson written down is repeated on every later run, stale notes mislead, and stored content can be planted by an attacker. Ask a supplier where the agent's memory lives, who can read it and how an entry is corrected.
Example
A procurement agent runs every night to chase late deliveries. Its notes file records that supplier K answers only through its portal, that orders under 2,000 euros are not escalated, and that 6 of last night's 37 reminders bounced. Tonight it reads the 40-line file first, uses the portal for K and retries only the 6. Without the file it would have sent all 37 again.
Most often confused with
Agent Memory vs. Memory (assistant feature)
The two use the same machinery, files and stores read back into the context, for different purposes. Assistant memory makes a product personal across conversations, and its hard questions are about privacy and consent. Agent memory keeps a piece of work coherent across steps and sessions, and its hard questions are about reliability: is the note still true, and who checks it? A product can have both.
Under the hood
A common vocabulary, borrowed from cognitive science and set out for language agents in the CoALA framework (Princeton, 2023), has four kinds. Working memory is the context window plus any scratchpad: what the agent holds now. Episodic memory records past runs: what was tried and how it ended. Semantic memory holds facts about the environment. Procedural memory holds how to do things: instructions, skills, code. Mechanisms: progress and notes files, to-do lists, the file system and version history, a memory tool with which the model reads and writes files that the application stores, vector stores searched for similar past episodes, and checkpoints of harness state for resuming a run. Each needs four operations: write (at milestones and before compaction), read (load at the start, retrieve on demand), consolidate and expire. Pitfalls: stale or wrong entries, unlimited growth, irrelevant retrieval that adds to context rot, leakage between users or tenants, and memory poisoning, in which injected content is saved and steers later runs. Keep memory scoped, readable by people and open to review.