
A multi-agent system is worth building when 3 things are true: the task splits into parts that do not depend on each other, each part needs more context or different tools than one agent can carry, and the answer is worth several times the tokens. When steps depend on each other, when every agent must agree on the same decisions, or when cost per task is tight, one agent does the job better. Everything else is detail hanging off those 3.
The usual wrong start is giving agents job titles: a planner, a researcher, a writer, a critic. That copies an org chart, and an org chart is not a reason to split a task. Anthropic reported in June 2025 that its multi-agent research system beat a single agent by 90.2 per cent and used about 15 times the tokens of a chat. Google researchers found in December 2025 that the same move cost 39 to 70 per cent of performance on sequential planning tasks. The failure is ordinary: 4 agents each make a reasonable decision about the same file, and the merged result works for none of them.
So the decision is about independence and coordination, not about how many agents you can draw. This post puts numbers on both.
Multi-agent systems are an old field with a new meaning. Research on autonomous agents that negotiate and cooperate goes back decades. When an engineering team says "multi-agent system" in 2026, they almost always mean something narrower: several large language model instances, each with its own context, prompt and tools, coordinated to complete 1 task.
The coordination is either code or a model. Anthropic's "Building effective agents", published in December 2024 by Erik S. and Barry Zhang, draws the line clearly. Workflows orchestrate models and tools through predefined code paths. Agents direct their own process and tool use. A multi-agent system can be either, or mix the two: code decides the shape, models decide the content.
The patterns have settled. The same post names 5 workflow patterns: prompt chaining, routing, parallelisation, orchestrator-workers and evaluator-optimiser. Orchestrator-workers is the pattern this post focuses on: a lead model breaks the task down, hands pieces to workers, and merges what comes back.
This is a different problem from building 1 good agent. How a single agent should use tools, retrieve from your documents and stay inside its permissions is covered in our posts on custom AI agent development and agentic RAG. Everything below assumes you already have 1 agent that works, and you are deciding whether to have several.
3 primary sources disagree on tone and agree on substance. Read them together and a clear rule emerges.
Anthropic's research system is the strongest case for splitting. In "How we built our multi-agent research system", published 13 June 2025 by Jeremy Hadfield and 5 colleagues, a lead agent on Claude Opus 4 directing subagents on Claude Sonnet 4 beat single-agent Claude Opus 4 by 90.2 per cent on Anthropic's internal research evaluation. The same post reports that agents use about 4 times the tokens of chat and multi-agent systems about 15 times, and that token usage alone explained 80 per cent of the performance variance on the BrowseComp benchmark. Research is breadth: many sources, read in parallel, condensed into 1 answer.
The same post names where it does not fit. Anthropic says domains where all agents must share the same context, or where agents depend heavily on each other, are a poor fit today, and notes that most coding tasks have fewer truly parallel parts than research.
Cognition's argument is the case against. In "Don't Build Multi-Agents", published 12 June 2025, Walden Yan sets out 2 principles: share full agent traces, not just individual messages, and remember that actions carry implicit decisions, so conflicting decisions produce bad results. That is the coding failure in 1 sentence.
Google's study puts numbers on the boundary. "Towards a Science of Scaling Agent Systems" by Yubin Kim and 19 co-authors at Google Research, Google DeepMind and MIT, first published 9 December 2025, tested 260 configurations across 6 benchmarks in its April 2026 version. On a parallel financial analysis task, centralised coordination improved results by 80.8 per cent. On a sequential planning task, every multi-agent variant got worse, by 39 to 70 per cent. Once a single agent already scored above about 45 per cent, adding agents gave negative returns.

The rule that survives all 3: split on independence, not on job titles. If 2 pieces of work can be done without either knowing what the other decided, they can be separate agents. If they cannot, they belong in 1 context.
This post groups multi-agent designs into 4 shapes, which real systems often stack. The difference that matters is where control and context live.
Shape | Who holds control | What each agent sees | Typical failure |
Orchestrator and workers | A lead agent or code, throughout | Its own scoped brief | The lead merges results that contradict each other |
Handoffs | Whichever agent was handed the conversation | The conversation so far | The task drifts as control moves |
Pipeline | Code, step by step | The previous step's output | An early mistake flows through unchecked |
Shared transcript | Nobody, by turn order | Everything every agent said | Cost and confusion grow with every turn |
Orchestrator and workers is the default for a reason. The OpenAI Agents SDK documentation calls this "agents as tools": a manager agent keeps control and calls specialists as tools. The lead sees summaries, not raw work, so its context stays small, and every worker starts clean.
Handoffs suit triage, not collaboration. In the same SDK, a handoff makes a specialist the active agent for the rest of the turn. LangGraph's swarm library works the same way and remembers which agent was last active. That is right for routing a customer to billing or support. It is wrong for a task where the first agent's judgement still matters after it hands over.
Pipelines are workflows wearing agent costumes. If the order of steps is fixed, write it as code and let each step be a model call. You gain determinism and lose nothing.
Shared transcripts are the shape to avoid by default. Every agent reads everything every other agent said. That feels like a meeting, and it costs like one.
Each worker costs roughly a full agent loop, because it is one. We wrote a small model of input tokens per task for 3 of the shapes, with every parameter exposed: token_model.py. A single agent runs 8 steps on a 3,000 token prompt, reading a 1,500 token tool result and writing 300 tokens each step. Each worker runs 6 steps on a 1,500 token brief and returns a 500 token summary. Shared transcript agents only talk, 400 tokens a message, 4 rounds each, and use no tools at all.

Orchestrator and workers grows in a straight line. About 36,500 input tokens per worker under these parameters, so 4 workers cost 2.0 times a single agent and 8 workers 4.0 times.
A shared transcript grows with the square of the turns. Going from 4 agents to 8 doubles the orchestrator's input tokens and triples the shared transcript's, from 96,000 to 294,400, even though these agents never call a tool. At 8 agents the talk-only transcript costs as much as 8 workers doing real work.
The model lands where Anthropic's numbers do. Anthropic's published multipliers imply a multi-agent system costs about 3.75 times a single agent: 15 times chat, divided by 4 times chat. Our model reaches 3.5 to 4 times a single agent at 7 to 8 workers. That is a sanity check on the parameters, not a measurement of your system. Replace them with numbers from your own traces before you trust either.
Model mixing changes the answer more than the shape does. Using the same parameters and Anthropic's list prices read on 6 October 2026, cost_model.py prices 1,000 tasks. Opus 5.5 is $4 per million input tokens and $20 per million output, Sonnet 5.5 is $2 and $10, and Haiku 4.5 is $1 and $5. No caching and no batch discount.
Setup | Per 1,000 tasks |
Single agent on Opus 5.5 | $346 |
Single agent on Sonnet 5.5 | $173 |
Lead on Opus 5.5, 4 workers on Opus 5.5 | $780 |
Lead on Opus 5.5, 4 workers on Sonnet 5.5 | $412 |
Lead on Opus 5.5, 4 workers on Haiku 4.5 | $228 |
Lead on Opus 5.5, 8 workers on Haiku 4.5 | $420 |
A multi-agent system can be cheaper than 1 large agent. A lead on Opus 5.5 with 4 workers on Haiku 4.5 costs $228 per 1,000 tasks, against $346 for a single agent on Opus 5.5, because 95 per cent of the tokens move to a model that costs a quarter as much. Whether Haiku workers are good enough for your subtasks is an evaluation question, not a pricing one.
The expensive mistake is uniform models. The same 4 workers on Opus 5.5 cost $780, 2.3 times the single agent, for the same shape. If you split a task, put the cheapest model that passes your evaluation on each worker and keep the large model for planning and merging.
3 things move these numbers. Prompt caching cuts the price of a repeated system prompt to a tenth or less of the input rate on these models. Anthropic notes that its Claude 4.7 and later tokenizer produces about 30 per cent more tokens for the same text, so token counts from older traces understate cost. And the number of steps per worker matters most of all: a worker that averages 9 steps instead of 6 costs about twice as much, because each step re-reads everything before it.
The boundary between agents is where the engineering goes. The Berkeley study below traces about three quarters of failures to system design and to misalignment between agents, not to the models. What decides whether the system works is what the orchestrator sends, what it accepts back, how long it waits, and what it does when a worker fails.
Write the contract down as code. Each worker gets a brief with its scope, the exact shape the orchestrator will accept, and its own token budget. Anything that breaks the contract is rejected, not repaired by hoping the lead model will notice. This is an excerpt from orchestrator.py, which runs as written with stub workers:
async def run_worker(name: str, brief: Brief, deadline_s: float) -> Result:
try:
data, used = await asyncio.wait_for(call_model(name, brief), deadline_s)
except asyncio.TimeoutError:
return Result(name, False, error=f"timed out after {deadline_s}s")
if used > brief.max_tokens:
return Result(name, False, tokens=used, error="over budget")
if not valid(data, brief.return_schema):
return Result(name, False, tokens=used, error="reply broke the contract")
return Result(name, True, data, used)
The output shows the point. We ran it with 3 stub workers, one of which is slow and one of which replies in the wrong shape. The orchestrator returned the 1 valid result, listed pricing as timed out after 1 second and reviews as having broken the contract, and marked the run "complete": false. Nothing was silently dropped and nothing off-contract reached the merge.
4 rules carry most of the weight. Give every worker a budget of its own, so 1 runaway worker cannot spend the whole run. Give every call a deadline, because a hung worker is otherwise indistinguishable from a slow one. Validate every reply against a schema before the lead model reads it. And return failures to the caller explicitly, so a partial answer is labelled partial rather than presented as complete.
Pass decisions, not just results. Cognition's first principle applies here. If a worker made a choice another worker depends on, such as a data format, a naming convention or an assumption about the user, that choice belongs in the contract. Otherwise 2 workers make 2 reasonable decisions and the merge inherits both.
The best data on failure comes from Berkeley. "Why Do Multi-Agent LLM Systems Fail?" by Mert Cemri, Melissa Z. Pan and 11 co-authors, published in March 2025 and accepted at NeurIPS 2025, annotated 1,642 execution traces from 7 open-source multi-agent frameworks. It identified 14 failure modes in 3 categories and reported failure rates between 41 and 86.7 per cent across the frameworks.
System design causes the largest share. In the October 2025 version of the paper, about 44 per cent of failures came from system design: disobeying the task or role specification, repeating steps, losing conversation history, and not knowing when to stop. About 32 per cent came from misalignment between agents, such as withholding information, ignoring another agent's input, or reasoning that does not match the action taken. About 24 per cent came from verification: stopping too early, checking incompletely or checking wrongly.

Small structural fixes moved the numbers. On ChatDev, clearer role specifications improved task success by 9.4 per cent. Adding a step that verifies the result against the high-level objective improved task success on the ProgramDev benchmark by 15.6 points, from 25.0 to 40.6 per cent in the paper's intervention table. Neither fix changed the model.
Independent agents amplify errors. The Google study measured error amplification by shape: 17.2 times for independent agents that do not check each other, against 4.4 times with a central coordinator. A lead that reviews worker output is not overhead. It is the error control.
The 2 protocols solve different problems, and you may need neither. Inside 1 system you control, plain function calls and a schema are simpler than any protocol.
MCP connects an agent to tools and data. Anthropic released the Model Context Protocol on 25 November 2024 and donated it on 9 December 2025 to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded by Anthropic, Block and OpenAI. The current specification is dated 28 July 2026. If your workers need the same tools, exposing those tools once over MCP saves writing each integration per agent.
A2A connects agents that belong to different owners. Google announced the Agent2Agent protocol on 9 April 2025, it became a Linux Foundation project on 23 June 2025, and it joined the Agentic AI Foundation on 27 August 2026. The specification reached version 1.0 in 2026. It is for an agent in your system talking to an agent in someone else's, where you cannot share code or memory.
Use A2A at the edge, not inside. Between your own orchestrator and your own workers, a protocol adds a network boundary, serialisation and another place for a contract to break. Reach for A2A when the other agent is a vendor's or a partner's.
Start with 1 agent and measure it. Log steps per task, tokens per step, and where the failures happen. You need these numbers to know whether splitting helps, and the Google result says that if your single agent already succeeds more than about 45 per cent of the time, more agents may make it worse.
Split the first piece that is genuinely independent. Usually that is a fan-out: searching several sources, checking several documents, or evaluating several options against the same criteria. Keep everything that requires a shared decision in the lead.
Give every worker the cheapest model that passes your evaluation. The cost table above is the argument. Test each worker's model on that worker's subtask, not on the whole task.
Make the lead verify, not just merge. The Berkeley result is that an explicit verification step against the original objective moved success more than any role change. The lead's last job is to check the merged answer against what was asked.
Trace everything from the first run. Every brief, every reply, every rejection, with timing and tokens. When 2 agents disagree in production, the trace is the only way to find out which handoff lost the context. Our AI agent development work starts with that tracing before it adds a second agent.
Decide what a partial answer looks like before you ship. Some workers will fail. Agree in advance whether the system returns a labelled partial result, retries once, or stops, and make the orchestrator do exactly that.
The short version. A multi-agent system helps when work splits into independent parts, and hurts when the parts depend on each other. Anthropic measured a 90.2 per cent gain on parallel research at about 15 times the tokens of chat; Google measured losses of 39 to 70 per cent on sequential planning. Orchestrator and workers grows linearly in cost, a shared transcript grows with the square of the turns, and mixing a large lead with small workers can cost less than 1 large agent. The model calls are the easy part. The handoff contract, the budget per worker, and the lead's final check against the original task are the system.
In current AI engineering, a multi-agent system is several large language model instances, each with its own context, prompt and tools, coordinated to complete 1 task. Coordination is done either by code, in a fixed workflow, or by a lead model that delegates work and merges results. The term is older than language models: it originally described any system of autonomous software agents that cooperate or negotiate.
Use several agents when the task splits into parts that do not depend on each other, such as researching many sources in parallel, and when the result justifies several times the tokens. Keep 1 agent when steps depend on each other or agents must agree on shared decisions. Google researchers found multi-agent setups improved a parallel task by 80.8 per cent but degraded sequential planning by 39 to 70 per cent.
Anthropic reported in June 2025 that its agents use about 4 times the tokens of chat and its multi-agent system about 15 times, implying roughly 3.75 times a single agent. Cost depends heavily on model choice per agent: at list prices on 6 October 2026, a lead on Claude Opus 5.5 with 4 workers on Haiku 4.5 can cost less than 1 agent on Opus 5.5 in a simple token model.
Orchestrator-workers is a multi-agent pattern in which a lead agent breaks a task into subtasks, sends each to a worker agent with its own context, and merges the results. Anthropic named it as 1 of 5 workflow patterns in "Building effective agents" in December 2024. It keeps the lead's context small, because the lead reads workers' summaries rather than their full work.
A study by Mert Cemri and colleagues at UC Berkeley, accepted at NeurIPS 2025, annotated 1,642 traces from 7 frameworks and found failure rates between 41 and 86.7 per cent. About 44 per cent of failures came from system design, about 32 per cent from misalignment between agents, and about 24 per cent from weak verification. Adding a step that checks the result against the original objective improved success by up to 15.6 per cent.
The Model Context Protocol connects an agent to tools and data sources, and was released by Anthropic in November 2024. The Agent2Agent protocol connects agents owned by different parties so they can delegate work to each other, and was announced by Google in April 2025. Both are now under the Agentic AI Foundation, part of the Linux Foundation. Inside a system you fully control, you need neither.
As few as the independent parts of the task require. Each worker in an orchestrator-workers design costs roughly a full agent loop, so cost grows in a straight line with the number of workers, and much faster in designs where every agent reads a shared transcript. Google's research also found that once a single agent succeeds more than about 45 per cent of the time, adding agents tends to reduce performance.

Vikas has around fifteen years of experience building software and now builds generative AI systems at Zyneto. His work covers retrieval augmented generation, agentic AI, knowledge graphs, AI memory, and the evaluation and guardrails that decide whether any of it is safe to put in front of customers. He has shipped enterprise copilots, document AI, chatbots and predictive analytics for e-commerce, fintech and marketing teams, and works day to day in Python, JavaScript and SQL. He follows multimodal models, business process automation and enterprise AI security closely, and mentors engineers moving into AI. He writes about architecture, inference cost and the failure modes that only show up at production scale.
Share your details and we will talk soon.
Be the first to access expert strategies, actionable tips, and the trends actually shaping the digital world. No fluff - just practical insights delivered straight to your inbox.
Dive into our blog and stay ahead of the curve with expert perspectives, future-ready trends, and tech tips written for decision-makers and doers alike.