In plain terms
A model on its own is an engine on a test bench: powerful, and going nowhere. The harness is the rest of the car: the transmission that turns power into movement, the steering, the brakes, the dashboard. When an agent reads a file, asks for approval before deleting something or stops after fifty steps, that is the harness at work. The model only ever proposes the next move; the harness makes the move happen, or refuses.
Why it matters
Two products built on the same model can differ widely in what they finish, what they cost and how safely they run, and the difference is the harness. When you buy an agent you are buying a harness as much as a model, and the controls a risk team asks about, such as permissions, sandboxing, approvals and logs, all live there. It is also where lock-in builds: swapping the model is usually a setting, swapping the harness is a migration. Two limits: a harness cannot rescue a weak model, and logic written to prop up last year's model can hold back this year's.
Example
A software company tests two coding agents that use the same model on 60 of its own maintenance tasks. One completes 31, the other 44. The second harness gives the model a code-search tool, clears stale tool output from the context, runs the tests after every edit and asks before deleting a file. Same model, 13 more tasks finished, and no deleted files to restore.
Most often confused with
Agent Harness vs. Model
People say “the model deleted the file” or “the model remembered”. The model wrote a request; the harness carried it out, and it was the harness that put the earlier notes back into the context. The distinction matters when something goes wrong and when products are compared: a benchmark score or a failure belongs to model and harness together, so ask which of the two would have to change.
Under the hood
A harness does six jobs. It assembles the prompt on every turn from the system prompt, tool definitions, history and memory files. It runs the agent loop and its stopping rules. It parses and executes tool calls, often in parallel, and handles errors and retries. It manages context through compaction, clearing of old tool results and subagents. It enforces safety: permission modes, approval prompts, sandboxed file and network access. It keeps state, so that a run can be paused, resumed and traced. Ready-made harnesses are offered as SDKs, among them the Claude Agent SDK and the OpenAI Agents SDK; LangGraph is a framework for building one's own; coding tools such as Claude Code and Codex are harnesses with a user interface. In evaluation work the same layer is called a scaffold. As models improve, harnesses tend to get thinner: steps that once needed code are left to the model, and the remaining code concentrates on tools, permissions and state.