Agents

Long-running Agent

Long-horizon agent

An AI agent that works on one task for hours or days, across many context windows and sessions, with little or no human input along the way.

ONE TASK · 14 HOURS · 9 CONTEXT WINDOWS · illustrativePERSONat milestonesAGENTwindow by windowFILESkept throughoutbriefcheck-in · 100check-in · 200review123456789↑ read↓ writeprogress file + commits · pending → done / blockedhour 0hour 14Each window starts by reading the notes and ends by writing them. By morning: 271 done, 29 blocked with reasons.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/long-running-agent

In plain terms

Most work with AI is a conversation: you ask, it answers, you ask again. A long-running agent is closer to a builder you leave in the house while you are away. You agree on the job, hand over the keys and come back to finished rooms, a list of what was done and a note about the one wall that needs your decision. No model can hold days of work in its head, so such an agent works in shifts, each starting from the notes the last one left.

Why it matters

It changes the unit of delegation from a question to a project: a code migration, a research review, a month-end reconciliation. The value per run is high, and so is everything else. A run costs hours of model use, and if it rests on a wrong assumption made in the first hour, all of it can be wasted. Nobody watches each step, so the safeguards have to be built in beforehand: a clear definition of done, tests the agent must pass, a budget cap, a sandbox, and milestones at which a person looks in.

Example

A retailer asks an agent to migrate 300 report scripts to a new data platform. It works for 14 hours overnight, across nine context windows. A progress file lists each script as pending, done or blocked; the agent commits after each one and marks it done only when a comparison test passes. Two check-ins reach the data lead, at 100 and at 200 scripts. By morning: 271 done, 29 blocked with reasons, and a bill of 180 dollars.

Most often confused with

Long-running Agent vs. Interactive agent session

Long-running AgentWorks unattended for hours; a person looks in at milestones
Interactive agent sessionA person reads each result and steers the next step

The same agent can be used either way; what changes is where the human is. In an interactive session mistakes are caught within minutes, because someone is reading along. In a long run the agent has to catch them itself, through tests and its own notes, and a person sees only checkpoints and the result. A batch job that runs all night is different again: it repeats one fixed step many times.

Under the hood

Four ingredients. Context management: compaction, clearing of old tool results and subagents keep each window usable. External state: a progress file, a task or test list with a status per item, and version-control commits as checkpoints; each new window begins by reading these and running the tests. Durable execution: the harness persists the run, so that it survives crashes, restarts and waits for approval, and can work in the background and notify on completion. Verification: tests, end-to-end checks or a separate reviewing agent, since the agent's own claim of completion is weak evidence. Failure modes: declaring the job finished early, attempting too much in one window and leaving work half-changed, goal drift, silent loops and runaway spend. On the trend, the research organisation METR tracks the length of software tasks, measured in human expert time, that frontier models complete with 50% success. In March 2025 it reported that this length had doubled about every seven months since 2019, and its later updates found a faster pace. The measure describes task difficulty; it makes no promise about unattended hours in a messy business setting.

Written by Mehmet Erkek · Last updated: