Subagents and the coordination tax
Splitting work across subagents buys exactly one thing - a smaller context per agent - and you pay for it twice, in duplicated tokens and in the answers that needed something sitting in another agent's window.
“Should I split this across subagents?” is the most common architecture question in agent engineering right now, and it is usually argued with adjectives. It has an arithmetic answer, and the arithmetic is not close.
Splitting does not make the work cheaper. It makes each context smaller, and those are completely different things.
What you pay
Every context pays the full briefing: the system prompt, the tool definitions, the task description. A parent with four children is five contexts, so that briefing is paid five times rather than once. Then each child pays again to report its findings back.
With 24 subtasks at 900 tokens each and a 4000-token briefing:
| Topology | Largest context | Total billed | Correct |
|---|---|---|---|
| One agent | 25,600 | 25,600 | 100% |
| Parent + 2 | 14,800 | 34,400 | 92.5% |
| Parent + 4 | 9,400 | 43,200 | 88.8% |
| Parent + 8 | 7,200 | 60,800 | 86.9% |
Read the two middle columns together, because that is the whole trade. Going to eight subagents cuts the largest context to 28% of what one agent needed, and costs 138% more tokens to do it.
Anthropic’s published numbers on their own multi-agent research system land in the same place from real production traffic: agents use roughly 4× the tokens of a chat interaction, and multi-agent systems roughly 15×.[2]articleHow we built our multi-agent research system
What you buy
Exactly one thing: headroom. A task that needs 60 subtasks cannot fit in a 40,000-token window as a single agent - it exhausts partway through and you get nothing. Split four ways, each child holds a quarter of the work and it completes.
That is a real and sufficient reason to reach for subagents. It is also the only one on this page, and it is worth being honest that “the work does not fit” is a narrower condition than “this task is complicated”.
The part that is easy to miss
Now raise the coupling.
A coupled subtask is one that needs something another subtask discovered. In a single
context that information is simply present; the agent scrolls up and there it is. Split
across k agents, it sits in a different window with probability (k-1)/k, and the
subagent has no way to know it is missing.
correctness ≈ 1 - coupling × (k-1) / k
At a modest 15% coupling and four subagents, that is roughly 11% of your subtasks answered from incomplete information. And the failure is silent - the subagent does not report that it lacked context, it answers confidently from what it had.
The empirical work agrees on where these systems break: a study of multi-agent LLM failures found the largest categories were not model errors at all but specification and inter-agent misalignment - agents proceeding on incomplete or inconsistent understanding of what the others were doing.[1]paperWhy Do Multi-Agent LLM Systems Fail?
Side by side
One agent
Everything in one context. Nothing is ever invisible to anything else.
Parent + 4 subagents
Five contexts, each paying the briefing again, each blind to the others.
| Measure | One agent | Parent + 4 subagents | Gap |
|---|---|---|---|
| Largest single context | 11200 | 5800 | 1.9× |
| Total tokens billed | 11200 | 27200 | 2.4× |
| Subtasks answered correctly | 1.00 | 0.88 | 1.1× |
System prompt, tool definitions, task description. Every context pays it again.
Share that need something another subtask found out.
The decision, as a procedure
Three steps, in order, and the order matters.
Compute the single-agent peak. Briefing plus subtasks times work per subtask. If it fits your window comfortably, stop - you have no problem that subagents solve, and splitting will cost you tokens and correctness for nothing.
If it does not fit, estimate the coupling. How many of these subtasks need what another one found? Independent work - search this list of files, summarise each of these documents - splits beautifully. Work where step nine depends on what step four turned up does not, and no amount of parent-level orchestration fixes it, because the parent only ever sees summaries.
Then pick the smallest fan-out that fits. Every extra agent buys less headroom than the last and costs another full briefing, another report back, and a larger share of coupled subtasks landing in the wrong window.[3]paperReAct: Synergizing Reasoning and Acting in Language Models
The dial
The line - when someone asks
Splitting a task across subagents does not make it cheaper, it makes each context smaller - and those are very different things. Every child pays the full briefing again and pays again to report back, so four subagents cost roughly 70% more tokens than doing the same work in one context while cutting the largest single context to about a third. The part that is easy to miss is correctness: a subtask that needs what another subtask discovered will find it sitting in a window it cannot see, and that failure is silent, because the subagent answers confidently from what it does have.
Recall
Loading…
Where are you with this?
Saved on this device. Sign in to keep it across devices.
Sources
Primary
- [1]
- [3]
Secondary
- [2]