In plain terms
At a hospital shift change the outgoing nurse briefs the incoming one before leaving: what each patient has had, what is due and what to watch for. From that moment the patients are the new nurse's responsibility. A handoff between agents, or from an agent to a human colleague, is the same moment: the work changes hands and a briefing travels with it. The quality of that briefing decides whether the receiver continues the task or starts it again.
Why it matters
Handoffs are where multi-step AI services most often break. A customer who is passed to a human and has to explain everything again has met a bad handoff; so has a specialist agent that repeats work or ignores a promise already made. Designing one means deciding what travels: the goal, what has been done, the facts gathered, the open question and the limits that apply. Too little, and the receiver guesses. Too much, and it inherits clutter, and personal data travels further than it needs to. Each handoff also adds delay and a point where accountability can blur.
Example
A bank's assistant takes a card dispute. It verifies the customer, collects the transaction details and finds that the amount, 1,240 euros, is above its 500-euro authority. It hands off to a human specialist with five lines: who the customer is, what is disputed, the evidence collected, what the customer was told, the decision needed. The specialist asks no repeat questions. Average handling time for such cases falls from 11 minutes to 4.
Most often confused with
Handoff vs. Subagent
Both pass work to another agent. A subagent is a delegate: it does one piece, reports back and ends, and the agent that called it stays in charge and answers to the user. After a handoff the first agent is out of the picture and the receiver deals with the user directly. The difference decides who needs the full history, and who is accountable for the outcome.
Under the hood
In several agent SDKs a handoff is presented to the model as a tool: calling it ends the current agent's turn and starts another agent with its own instructions and tools. In the OpenAI Agents SDK, for example, the receiving agent sees the whole conversation by default, and an input filter can trim it. Context is passed in one of two ways: the full transcript, which is complete, long and may carry clutter or injected text, or a structured brief, which is compact and only as good as its writer. Typical patterns: a triage agent routing to specialists, the stages of a pipeline, and escalation to a person. What should travel: goal and success criteria, work done, references to files or records, decisions with their reasons, constraints and a trace ID. Permissions should be granted to the receiver afresh and never inherited through text. Pitfalls: agents passing a task back and forth, constraints lost in a summary, duplicated work. Between agents from different vendors, the A2A protocol plays this role.