"Multi-agent" is the most over-prescribed idea in AI engineering right now. Teams split a simple task across a planner, a researcher, a writer and a critic, then spend weeks debugging the conversations between them. Sometimes that architecture is exactly right. More often a single agent, or even a fixed workflow, would have been faster, cheaper and easier to fix.
This guide gives you a way to decide, based on the clearest public evidence available: Anthropic's own write-ups on building agents and on the multi-agent research system behind Claude's Research feature. It ends with a small orchestrator in Python, tested, that shows what you actually gain and what you give up.
Start with the simplest thing that works
Anthropic's guide "Building effective agents" (December 2024) separates two kinds of systems. Workflows are "systems where LLMs and tools are orchestrated through predefined code paths". Agents are "systems where LLMs dynamically direct their own processes and tool usage". Its central advice is blunt: "We recommend finding the simplest solution possible, and only increasing complexity when needed."
That gives you a ladder, and you should climb it one rung at a time:
- A single model call with good instructions and examples. Most "agent" ideas end here.
- A workflow, when the steps are known in advance. The same guide describes the common patterns: prompt chaining (fixed steps, each using the previous output), routing (classify the input and send it to a specialized path), parallelization (split independent parts, or run several attempts and vote), and evaluator-optimizer (one model drafts, another critiques, repeat).
- A single agent, when the path can't be fixed in advance and the model must decide which tools to call.
- Multiple agents, usually as orchestrator-workers: a lead agent breaks the job down and delegates to worker agents. The guide recommends it for "complex tasks where you can't predict the subtasks needed".
What the evidence says about multi-agent systems
In June 2025 Anthropic published how it built Claude's multi-agent research system. The numbers are the most useful public data on this question, so it is worth reading them carefully:
- It can be much better. A system with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents "outperformed single-agent Claude Opus 4 by 90.2%" on Anthropic's internal research evaluation.
- It is much more expensive. "Agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats."
- The two facts are linked. In their analysis of the BrowseComp benchmark, "token usage by itself explains 80% of the variance" in performance. Multi-agent systems win largely because they let you spend more tokens usefully, in parallel, across separate context windows.
The same write-up is just as clear about where the approach fits. It works best for "valuable tasks that involve heavy parallelization, information that exceeds single context windows, and interfacing with numerous complex tools". It is a poor fit for "domains that require all agents to share the same context or involve many dependencies between agents", and the authors note that "most coding tasks involve fewer truly parallelizable tasks than research".
Our position: a four-question test
Go multi-agent only if you can answer "yes" to all four of these. If any answer is "no", a single agent or a workflow will serve you better.
- Is the work breadth-first? Can it split into parts that don't need each other's results, such as researching five vendors, reading ten documents, or checking six markets? If step 3 needs step 2's output, parallel workers just wait for each other.
- Is there more material than one context window holds well? If a single agent can read everything comfortably, splitting it only adds handoffs.
- Is the task valuable enough to pay several times more? At roughly 15× the tokens of a chat, a multi-agent run must be worth it. A due-diligence report, yes; a customer-support reply, almost never.
- Can you evaluate it? Multi-agent bugs hide in the handoffs. Without an eval suite you won't know whether the extra agents helped. (Our guide to testing agents with pass^k covers how.)
See the trade-off in code
The script below runs the same research job, covering five independent subtasks, in two architectures. The workers are stand-ins with a fixed delay, so the run is reproducible without an API key; in a real system you replace call_model with a model call and keep the rest. Create orchestrator.py:
"""Single agent vs orchestrator-workers: the same research job, two architectures.
Workers are stand-ins with a fixed delay so the run is reproducible. Swap
`call_model` for a real model call; the orchestration code stays the same.
"""
import asyncio
import time
SYSTEM_PROMPT_TOKENS = 1_500 # instructions + tool definitions every agent reads
RESULT_TOKENS = 2_000 # raw search results for one subtask
SUMMARY_TOKENS = 200 # what a worker hands back to the lead
STEP_SECONDS = 0.3 # stand-in latency of one subtask (scaled down)
async def call_model(subtask: str, fail: bool = False) -> str:
"""Stand-in for one agent researching one subtask."""
await asyncio.sleep(STEP_SECONDS)
if fail:
raise TimeoutError(f"search tool timed out on {subtask!r}")
return f"key findings on {subtask}"
async def single_agent(subtasks: list[str]) -> dict:
"""One agent, one context window: raw results pile up as it goes."""
start, context, peak = time.perf_counter(), SYSTEM_PROMPT_TOKENS, 0
for s in subtasks:
await call_model(s)
context += RESULT_TOKENS
peak = max(peak, context)
return {"seconds": time.perf_counter() - start, "peak_context": peak}
async def orchestrator(subtasks: list[str], failing: frozenset = frozenset()) -> dict:
"""A lead agent fans out to workers in parallel; each worker has a clean context."""
start = time.perf_counter()
jobs = [call_model(s, fail=s in failing) for s in subtasks]
results = await asyncio.gather(*jobs, return_exceptions=True) # one failure must not sink the rest
found = [r for r in results if not isinstance(r, Exception)]
gaps = [f"{s}: {r}" for s, r in zip(subtasks, results) if isinstance(r, Exception)]
worker_peak = SYSTEM_PROMPT_TOKENS + RESULT_TOKENS
lead_peak = SYSTEM_PROMPT_TOKENS + SUMMARY_TOKENS * len(found) # lead reads summaries only
return {"seconds": time.perf_counter() - start,
"peak_context": max(worker_peak, lead_peak), "gaps": gaps}
async def main():
subtasks = ["pricing", "API limits", "data residency", "Arabic support", "SLA terms"]
single = await single_agent(subtasks)
multi = await orchestrator(subtasks)
print(f"{'architecture':<22}{'seconds':>9}{'largest context':>17}")
print(f"{'single agent':<22}{single['seconds']:>9.1f}{single['peak_context']:>17,}")
print(f"{'orchestrator-workers':<22}{multi['seconds']:>9.1f}{multi['peak_context']:>17,}")
partial = await orchestrator(subtasks, failing=frozenset({"SLA terms"}))
print("\nWith one worker failing, the lead still answers and reports the gap:")
for g in partial["gaps"]:
print(" missing ->", g)
if __name__ == "__main__":
asyncio.run(main())
Three details make this a real orchestrator rather than a loop:
asyncio.gatherruns the workers at the same time. The lead doesn't wait for one subtask before starting the next.- Each worker gets a clean context. It sees only the instructions and its own subtask, not the other four workers' raw results. That is the "separate context windows" advantage in Anthropic's write-up.
return_exceptions=Trueisolates failures. If one worker's tool times out, the others still return, and the lead reports the gap instead of crashing. Anthropic's write-up stresses that in agent systems "minor system failures can be catastrophic", so this matters more than it looks.
What the run shows, and what it doesn't
Running python orchestrator.py printed:
architecture seconds largest context
single agent 1.5 11,500
orchestrator-workers 0.3 3,500
With one worker failing, the lead still answers and reports the gap:
missing -> SLA terms: search tool timed out on 'SLA terms'
Read it carefully, because it shows both sides of the trade:
- Speed: five times faster here. The five independent subtasks ran together, so the job took one step's time instead of five. That only happens because the subtasks were independent; with a chain of dependent steps, the two numbers would be the same.
- Context: no agent held more than 3,500 tokens. The single agent ended with 11,500 tokens in one window (1,500 of instructions plus 5 × 2,000 of raw results), and that grows with every subtask. The lead in the orchestrator reads only short summaries. This is why multi-agent systems handle "information that exceeds single context windows".
- Failure: the job still finished. One worker failed, and the output names exactly what is missing. That is the behavior you want from a lead agent: a partial, honest answer instead of nothing.
What the run does not show is total cost. Our stand-in workers do a fixed amount of work, but real subagents each run their own search loops and re-read their own instructions, which is why Anthropic measured about 15× the tokens of a chat. Budget for the higher bill; don't expect the parallelism to pay for itself.
If you do build one: lessons from production
- Teach the lead to scale effort to the question. Anthropic embedded explicit guidance in its lead agent's prompt: simple fact-finding needs "just 1 agent with 3-10 tool calls", direct comparisons "might need 2-4 subagents with 10-15 calls each", and complex research "might use more than 10 subagents". Without such rules, leads tend to spawn too many workers for easy questions.
- Write worker briefs like task tickets. Each worker needs an objective, an output format, which tools to use, and clear boundaries, so two workers don't research the same thing.
- Return summaries, not raw dumps. The lead's context is the bottleneck. Workers should hand back condensed findings with sources, as in the example.
- Plan for long-running failures. Agent runs are stateful: a crash halfway through can't simply restart from zero. Save progress, retry tools, and let the agent adapt when a tool fails.
- Deploy carefully. Anthropic uses "rainbow deployments", gradually shifting traffic from old to new versions so that changes don't break agents already mid-run.
Bottom line
A multi-agent system is a tool for one specific shape of problem: broad, parallel, high-value work that overflows a single context window. For that shape the gains are large and well documented. For everything else, the extra agents add cost, latency between handoffs, and new places to fail. Climb the ladder one rung at a time, measure each step with evals, and add agents only when the evidence says the next rung pays for itself.
The building blocks covered in our earlier guides still apply inside every worker: tools exposed over MCP and memory that survives long runs.
Sources: Anthropic Engineering, "Building effective agents" (December 19, 2024) and "How we built our multi-agent research system" (June 13, 2025). Code tested on 2026-10-05 with Python 3.13; the workers are stand-ins with fixed delays and token sizes, so the speed and context numbers illustrate the architecture and are not measurements of any real model.