Agents

Agent Harness

Agent scaffold

The software around a language model that turns it into a working agent: it runs the loop, executes tool calls, manages context and memory, enforces permissions and decides when to stop.

AGENT HARNESS · everything around the modelAgent loopruns the turnsContext managementcompacts, clears, delegatesMemory and statenotes, checkpoints, resumePrompt assemblywhat the model sees each turnModelproposes the next stepTool executionruns what the model requestsPermissionsallow, ask or denySandboxlimits files and networkStopping rulessteps, budget, timeKeep the model and change the harness, and you have a different agent.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/agent-harness

In plain terms

A model on its own is an engine on a test bench: powerful, and going nowhere. The harness is the rest of the car: the transmission that turns power into movement, the steering, the brakes, the dashboard. When an agent reads a file, asks for approval before deleting something or stops after fifty steps, that is the harness at work. The model only ever proposes the next move; the harness makes the move happen, or refuses.

Why it matters

Two products built on the same model can differ widely in what they finish, what they cost and how safely they run, and the difference is the harness. When you buy an agent you are buying a harness as much as a model, and the controls a risk team asks about, such as permissions, sandboxing, approvals and logs, all live there. It is also where lock-in builds: swapping the model is usually a setting, swapping the harness is a migration. Two limits: a harness cannot rescue a weak model, and logic written to prop up last year's model can hold back this year's.

Example

A software company tests two coding agents that use the same model on 60 of its own maintenance tasks. One completes 31, the other 44. The second harness gives the model a code-search tool, clears stale tool output from the context, runs the tests after every edit and asks before deleting a file. Same model, 13 more tasks finished, and no deleted files to restore.

Most often confused with

Agent Harness vs. Model

Agent HarnessRuns the loop, executes the tools, enforces the limits
ModelReads the context and proposes the next step

People say “the model deleted the file” or “the model remembered”. The model wrote a request; the harness carried it out, and it was the harness that put the earlier notes back into the context. The distinction matters when something goes wrong and when products are compared: a benchmark score or a failure belongs to model and harness together, so ask which of the two would have to change.

Under the hood

A harness does six jobs. It assembles the prompt on every turn from the system prompt, tool definitions, history and memory files. It runs the agent loop and its stopping rules. It parses and executes tool calls, often in parallel, and handles errors and retries. It manages context through compaction, clearing of old tool results and subagents. It enforces safety: permission modes, approval prompts, sandboxed file and network access. It keeps state, so that a run can be paused, resumed and traced. Ready-made harnesses are offered as SDKs, among them the Claude Agent SDK and the OpenAI Agents SDK; LangGraph is a framework for building one's own; coding tools such as Claude Code and Codex are harnesses with a user interface. In evaluation work the same layer is called a scaffold. As models improve, harnesses tend to get thinner: steps that once needed code are left to the model, and the remaining code concentrates on tools, permissions and state.

Written by Mehmet Erkek · Last updated: