In plain terms
One agent doing everything is a generalist with an overflowing desk. A multi-agent system is a small team: a lead who splits the work, specialists who each handle a part, and someone who assembles the results. Each member keeps a clear head because each sees only its own part.
Why it matters
For broad tasks that split into independent parts, a team of agents can be faster and more thorough than one. It also costs several times more in tokens and is harder to debug. The sensible default is a single agent; add more when the task is clearly bigger than one context window or can truly run in parallel.
Example
A market study: a lead agent breaks the question into five sub-questions, five research agents investigate in parallel, each returns a short findings note, and the lead writes the combined report with sources.
Most often confused with
Multi-agent System vs. Agentic Workflow
A workflow with five model calls is not a multi-agent system. What makes it multi-agent is that each participant is an agent: it runs its own loop and chooses its own steps. If the steps are fixed in code, it is a workflow, however many prompts it has.
Under the hood
Common topologies: orchestrator–workers (a lead delegates to subagents and synthesises), pipelines with handoffs, and peer networks. The main benefit is context isolation: each agent spends its window on one sub-problem, and work can run in parallel. The main costs are token use, coordination failures (duplicated or missing work when briefs are vague) and lost detail when results are summarised. Tasks with tightly coupled parts, where every step depends on shared state, rarely benefit.