Agents & Tool Integration · 6 min
Multi-Agent Orchestration: When It Helps, When It's Overkill
How to tell real task decomposition from cargo-culted agent sprawl, using honest numbers on cost, latency, and coordination failure.
You've built a single agent: a model in a loop with tools, running until the task is done or a budget runs out. It works, mostly, but on hard tasks it thrashes. The tempting next move is to split it into a team: a planner, some specialists, maybe a critic. Sometimes that's the right call. More often it's expensive theater, and this lesson is about telling the two apart.
The mental model: parallelism, not org charts
A multi-agent system is not a company. The useful abstraction is far narrower. Ask one question: can this task be cut into subtasks whose intermediate context does not need to be shared while they run?
If yes, separate agents let you parallelize the work and give each subtask its own clean context window. No cross-contamination, no single 200k-token buffer holding six half-finished threads at once. If no, if the subtasks constantly need to see each other's decisions, then you don't have parallel work. You have one sequential task wearing a costume, and splitting it will hurt.
Anthropic's research system is the canonical "yes" case. A lead agent decomposes a query, spins up 3–5 subagents in parallel, each searches an independent branch, and the lead synthesizes. It beat single-agent Opus 4 by 90.2% on their internal research eval, but specifically for breadth-first queries that "involve pursuing multiple independent directions simultaneously." That qualifier is the whole game.
Three patterns, honestly
Orchestrator / worker. A lead agent plans and delegates, workers execute in isolated contexts, the lead synthesizes. This is the pattern that actually earns its cost, and only when the branches are genuinely independent.
┌─────────────┐
query ───────▶│ orchestrator│ plans, splits, synthesizes
└──────┬──────┘
┌────────────┼────────────┐
▼ ▼ ▼
[worker A] [worker B] [worker C] ← isolated contexts,
search search search run in parallel
└────────────┼────────────┘
▼
final synthesis (+ separate citation pass)Specialist roles. Fixed agents, each with a tuned prompt and a narrow toolset: a SQL writer, a security reviewer, a summarizer. What this actually buys you is smaller, auditable tool surfaces per role, more than it buys raw quality. A single agent holding all the tools with a good system prompt often matches it, because the model already routes between tools internally.
Debate / critique. N agents answer, then critique each other over several rounds, converging on a consensus. It's seductive, and it's the shakiest of the three. Du et al. (2023) showed multi-agent debate improving arithmetic and factuality as you add agents and rounds. But a 2025 evaluation, Stop Overvaluing Multi-Agent Debate, ran five representative debate methods across nine benchmarks and found they often fail to beat single-agent baselines like chain-of-thought and self-consistency, even while burning far more inference compute. The version that survives scrutiny is a generate-and-verify split: one agent drafts, one checks against a rubric. That's asymmetric critique, not a debate club.
The numbers you're actually signing up for
Memorize Anthropic's measurement. Single agents use about 4× the tokens of a chat interaction; multi-agent systems use about 15×. In their eval, token usage alone explained 80% of the variance in performance, with tool-call count and model choice as the other two factors. That cuts both ways. A large share of the "multi-agent is better" result is just more compute thrown at the problem.
So before reaching for a swarm, ask whether one agent with a bigger budget gets you most of the way. Or self-consistency: sample the single agent k times and vote. Either buys you a chunk of the benefit for a fraction of the coordination risk.
In Anthropic's Research system, multi-agent runs consumed roughly 15× the tokens of normal chat. That's not a universal multiplier — your architecture may be far below or above it — but recursive delegation and duplicated context can make the cost explode surprisingly quickly. A worker that recursively spawns more workers, or a tool that returns an oversized payload N agents each re-read, compounds it fast. And cost isn't only dollars. Workers run in parallel, but the orchestrator's plan and its final synthesis are serial bottlenecks. You pay fan-out latency (the slowest worker) plus two sequential model passes. For an interactive feature this often feels slower than a single streaming agent even when total throughput is higher.
Builder: Instrument before you distribute. Log per-agent token counts and wall-clock time. If one agent burns 70% of the budget, you have a tooling or prompt problem, not an architecture problem. A second agent won't fix it. It'll double it.
How coordination actually fails
Cognition's Don't Build Multi-Agents is the essential counterweight, and its argument is concrete. Two principles: share full agent traces, not just individual messages, and actions carry implicit decisions, so conflicting decisions carry bad results. Their example: asked to build a Flappy Bird clone, one subagent renders a Super Mario Bros–style background while another builds a bird that doesn't match and moves wrong. Neither output is wrong in isolation. The assembly is incoherent because the assumptions were never shared. Naive parallelism manufactures this class of bug.
You'll meet these failure modes in the wild:
- Assumption drift. Workers silently make incompatible choices (the Flappy Bird problem).
- Lost context. Passing summaries instead of full traces drops the decision that made a later step make sense.
- Runaway spawning. Early versions of Anthropic's system launched 50 subagents for a simple query and scoured the web for nonexistent sources. Delegation is itself a skill models are still weak at.
- Duplicated work. Vague delegation sends two workers on the same search, and the synthesizer double-counts the result as corroboration.
The security angle: coordination is attack surface
Each agent boundary is a trust boundary, and most orchestration code treats worker output as clean data. It isn't. If worker A retrieves a poisoned web page carrying an injected instruction, and the orchestrator ingests A's output verbatim, the injection has propagated across an agent hop. Now it can shape how other workers are tasked. A multi-agent topology turns a single indirect prompt injection into a lateral-movement problem.
poisoned page → worker A → "instruction" rides in output
│
▼
orchestrator ── re-delegates ──▶ worker B, C ← blast radiusThe mitigations are old news in new clothes. Treat every inter-agent message as untrusted input. Keep per-worker tool permissions least-privilege (a search worker never needs write or shell access). Put the irreversible actions behind one narrow, audited agent instead of letting every specialist hold dangerous tools. A wide fan-out with uniform permissions means your blast radius is the union of everything any agent can do.
Defender: Log the full trace, every inter-agent message and every tool call with its arguments, as one correlated session rather than per-agent silos. When a run goes wrong you need to see which hop the bad decision entered. That's also what Cognition's "share full traces" principle demands, so reliability and auditability fall out of the same plumbing.
Researcher: Hold compute constant. A multi-agent result that uses 15× the tokens has to be compared against a single agent given 15× the budget, or against 15 self-consistency samples, not against one cheap pass. Most published "multi-agent wins" quietly skip this control, which is exactly why the 80%-variance-from-tokens finding matters.
A decision rule
Default to the simplest loop that works: one well-tooled agent, measured. Reach for multiple agents only when all three conditions hold. The work splits into genuinely independent branches. Each branch would otherwise pollute a shared context. And the task's value justifies a 15× cost with two serial coordination passes. Research, broad literature review, and large-codebase reconnaissance fit. Tightly interdependent work does not: most coding, and anything where every step depends on the last, is better served by a single agent with context compression than by a committee. If you can't name which independent subtasks you're parallelizing, you don't need multiple agents. You need a better prompt and better tools. (See the RAG-vs-tools lesson in this module for where retrieval belongs inside that single loop.)
Sources
- Anthropic, "How we built our multi-agent research system" (engineering blog, 2025). Source for the 4×/15× token figures, the 90.2% improvement, 3–5 subagents, the 80%-of-variance and 50-subagent findings. — https://www.anthropic.com/engineering/multi-agent-research-system
- Cognition (Walden Yan), "Don't Build Multi-Agents" (2025). Two principles, Flappy Bird example, context engineering. — https://cognition.com/blog/dont-build-multi-agents
- Du, Li, Torralba, Tenenbaum, Mordatch, "Improving Factuality and Reasoning in Language Models through Multiagent Debate," arXiv:2305.14325 (2023).
- Zhang et al., "Stop Overvaluing Multi-Agent Debate: We Must Rethink Evaluation and Embrace Model Heterogeneity," arXiv:2502.08788 (2025). MAD often fails to beat CoT and self-consistency at matched or higher compute.