In plain terms
A manager who reads every document personally runs out of attention. Instead she sends an assistant to go through the archive and come back with a one-page summary. A subagent is that assistant: it does the digging in its own workspace, so the main agent's memory stays free for the bigger picture.
Why it matters
Subagents are the main technique for keeping long tasks on track, because an agent whose context fills with raw search results gets slower, costlier and less accurate. The trade-off is visibility: you see the summary, not the forty steps behind it, so the brief you give and the report you ask for matter.
Example
A coding agent has to find every place a function is used across a large codebase. It starts a subagent with that one job. The subagent reads two hundred files and returns a list of twelve locations; the main agent continues with the list and none of the clutter.
Most often confused with
Subagent vs. Tool
Calling a tool is like pressing a button: one fixed operation, one result. Starting a subagent is like delegating: it has a model, takes many steps and uses judgment. To the main agent both look similar, a request out and a result back, which is why subagents are often exposed as a tool.
Under the hood
A subagent gets a fresh context window containing only its brief, usually with its own system prompt, tool subset and sometimes a cheaper model. Only its final message returns to the caller, so intermediate tool output never enters the parent context. Good briefs are self-contained: goal, what done looks like, constraints, and the format of the answer. Several subagents can run in parallel. Costs: total tokens rise, the parent cannot see intermediate reasoning, and a subagent inherits no conversation history unless it is passed in.