Prompting & context

Memory

Any mechanism that lets an AI system carry information from one conversation or session to another, since the model itself retains nothing between requests.

Context windowshort-term memoryModelstateless by itselfMemory storelong-term memorywritereadgone when the session endskept across sessionsA model remembers nothing between conversations; “remembering” is old information being put back into context.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/memory

In plain terms

A language model has no memory of its own. Each request starts blank, and anything it seems to recall about you was placed in front of it again by the application. “Memory” is that plumbing: the system writes down things worth keeping, such as your preferences, past decisions and the project's history, and slips the relevant ones back into the context next time.

Why it matters

Memory is what turns a clever tool you must re-brief every morning into an assistant that knows your business. It also raises questions a company has to answer on purpose: what is stored, for how long, who can read or delete it, and whether one customer's information could ever surface in another's session. Memory features deserve the same scrutiny as any system that holds personal data.

Example

A project assistant is told in March that the client prefers weekly summaries on Fridays and dislikes slide decks. In June, with no reminder, it drafts the Friday summary as a short document. It did not “remember”: a note saved in March was retrieved and placed in the context when the June conversation began.

Most often confused with

Memory vs. Context Window

MemoryStored outside the model; persists across sessions
Context WindowInside the current request; gone when it ends

The context window is short-term working space, limited in size and wiped after each session. Memory is long-term storage outside the model. They work together: memory holds far more than fits in a window, and only the pieces relevant to the moment are loaded into the context.

Under the hood

Implementations range from simple to elaborate: a file of notes loaded at session start (an instructions or memory file); a memory tool the model uses to read and write files itself; vector stores of past interactions searched by similarity; and structured profiles or knowledge graphs. Design questions: what triggers a write, how memories are updated or retired when facts change, how retrieval avoids pulling stale or irrelevant items, and how users can inspect and delete entries. Risks: stale memories treated as current, sensitive data retained without consent, and memory poisoning, in which injected content is saved and then influences later sessions.

Written by Mehmet Erkek · Last updated: