Home › Agents & Tool Use › Swarm / handoff
🕹️ · Build

Swarm / handoff

Peer agents route work to each other by calling explicit handoff functions, with no central controller.

In one line

A swarm is a set of peer agents that transfer control to one another through explicit handoff rules, so routing emerges locally instead of from a central orchestrator.

ConceptWhat it is

A swarm is a decentralised multi-agent pattern: a set of peer agents, each with its own instructions and tools, that pass control to one another by calling an explicit handoff. There is no central orchestrator deciding who runs next — an agent that recognises a request belongs to a different specialty simply hands the live conversation over, and the receiving agent takes the wheel.

The pattern exists to make skill-based routing lightweight. Instead of encoding every route in one large planner, each agent only needs to know its own job and which peers it may hand off to. Routing then emerges from these local rules, which keeps individual agents small and focused — at the cost of a system whose overall path is harder to predict and trace.

How it worksThe mechanics

Each agent is defined with a prompt, a toolset, and a set of handoff targets exposed to the model as callable functions; the model reads the incoming message and either answers using its own tools or calls a handoff, which transfers control plus the accumulated conversation state to the named peer, who resumes from that context and repeats the loop until some agent replies directly to the user.

At a glanceSee it

Swarm / handoff diagram
Swarm / handoff diagram 1

Under the hood a handoff is just a tool call the runtime intercepts — it repoints the active agent, swaps the prompt and tools, and carries the shared history forward.

Swarm / handoff diagram 2

Swarms fail in a few characteristic ways — endless ping-pong, ballooning context, and orphaned requests — each answered by a specific guardrail.

When to use itWhere it fits

  • Support-style routing where a request must reach the right specialist skill — billing, returns, technical — and the categories are known.
  • You want to add or retire a skill by editing one agent's handoff list rather than rewiring a central router.
  • Sub-tasks are served better by focused, narrowly-prompted agents than by one agent juggling every domain.
  • Conversation state should follow the user across skills without rebuilding context at each hop.

When NOT to use itLimits & anti-patterns

  • The workflow has a fixed, auditable sequence of steps — a planner-executor or explicit graph is easier to reason about.
  • You need tight guarantees over which agent handled what, since emergent routing is hard to trace and reproduce.
  • Latency is critical, because each handoff is another model round-trip.
  • A single well-prompted agent already covers the task, so extra agents add cost with no quality gain.

Trade-offsAdvantages & costs

Advantages
  • Decentralised — there is no central orchestrator to design, and each agent stays small and independently testable.
  • Easy to extend: adding a skill means adding an agent and one handoff edge, not touching a router.
  • Handoffs carry the conversation state, so the user is not made to repeat themselves across skills.
  • Lightweight to stand up compared with a full orchestrator-worker or graph.
Trade-offs & costs
  • Behaviour is emergent, so the end-to-end path is harder to trace, debug, and reproduce.
  • Agents can ping-pong or hand off in loops without a step budget or loop guard.
  • Every handoff is another model call, so tokens and latency grow with the number of hops.
  • Passing full history at each handoff can bloat context and cost unless it is filtered.

ExampleIn the real world

A customer-support assistant starts every conversation at a triage agent whose only job is to read the opening message and hand off. A shipping question triggers a handoff to a logistics agent with order-lookup tools; if that agent discovers the real issue is a refund, it hands off again to a billing agent that can process one. The full thread travels with each handoff, so the customer never re-explains the problem, and the team can add a new capability — say, a warranty agent — by defining one more agent and listing it as a handoff target, with no change to the others.

ToolsHow to implement it

  • OpenAI Swarmthe experimental, educational library that popularised returning an agent from a function to transfer control.
  • OpenAI Agents SDKits production successor, where handoffs are a first-class primitive represented as tools the model can call.
  • LangGraphmodels peer handoffs as edges over shared state; the langgraph-swarm extension packages the pattern directly.
  • CrewAIsupports delegation between role-based agents, a related handoff-style flow.

Cost & effortWhat it takes

Cost scales with the number of handoffs: every hop is a fresh model call, and because the conversation history usually travels with control, later agents read a longer, more expensive context than the first. Engineering effort is modest to start — defining a few agents and their handoff lists is quick — but rises once you add the loop guards, step budgets, and tracing needed to keep emergent routing safe and debuggable in production.

A living map of modern AI — kept current every morning