yield point
Agentsintermediateupdated 2026-08-23

Subagents and the coordination tax

Splitting work across subagents buys exactly one thing - a smaller context per agent - and you pay for it twice, in duplicated tokens and in the answers that needed something sitting in another agent's window.

“Should I split this across subagents?” is the most common architecture question in agent engineering right now, and it is usually argued with adjectives. It has an arithmetic answer, and the arithmetic is not close.

Splitting does not make the work cheaper. It makes each context smaller, and those are completely different things.

Run it as one agent, then switch the topology to subagents and watch the largest context collapse while the total bill climbs. Then raise the coupling and see what the split quietly costs you in answers.

Largest context4,000 / 64,000
Answered correctlyNeeded something it could not see
Topology
one agent
Subtasks
0 / 24
Largest context
4,000 tok
Total billed
4,000 tok
Correct
100.0%
tick 0 / 100
Break it

System prompt, tool definitions, task description. Every context pays it again.

Share that need something another subtask found out.

Safety properties
  • ✓
    No single context exceeds the budget

    held on every tick so far

Measured vs theory
MetricSimulatedFormulaError
peakContext40002560084.4%
tokensBilled40002560084.4%

What you pay

Every context pays the full briefing: the system prompt, the tool definitions, the task description. A parent with four children is five contexts, so that briefing is paid five times rather than once. Then each child pays again to report its findings back.

With 24 subtasks at 900 tokens each and a 4000-token briefing:

Topology Largest context Total billed Correct
One agent 25,600 25,600 100%
Parent + 2 14,800 34,400 92.5%
Parent + 4 9,400 43,200 88.8%
Parent + 8 7,200 60,800 86.9%

Read the two middle columns together, because that is the whole trade. Going to eight subagents cuts the largest context to 28% of what one agent needed, and costs 138% more tokens to do it.

Anthropic’s published numbers on their own multi-agent research system land in the same place from real production traffic: agents use roughly 4× the tokens of a chat interaction, and multi-agent systems roughly 15×.[2]articleHow we built our multi-agent research systemAnthropic, 2025

What you buy

Exactly one thing: headroom. A task that needs 60 subtasks cannot fit in a 40,000-token window as a single agent - it exhausts partway through and you get nothing. Split four ways, each child holds a quarter of the work and it completes.

That is a real and sufficient reason to reach for subagents. It is also the only one on this page, and it is worth being honest that “the work does not fit” is a narrower condition than “this task is complicated”.

The part that is easy to miss

Now raise the coupling.

A coupled subtask is one that needs something another subtask discovered. In a single context that information is simply present; the agent scrolls up and there it is. Split across k agents, it sits in a different window with probability (k-1)/k, and the subagent has no way to know it is missing.

correctness ≈ 1 - coupling × (k-1) / k

At a modest 15% coupling and four subagents, that is roughly 11% of your subtasks answered from incomplete information. And the failure is silent - the subagent does not report that it lacked context, it answers confidently from what it had.

The empirical work agrees on where these systems break: a study of multi-agent LLM failures found the largest categories were not model errors at all but specification and inter-agent misalignment - agents proceeding on incomplete or inconsistent understanding of what the others were doing.[1]paperWhy Do Multi-Agent LLM Systems Fail?Cemri, M. et al., 2025

Side by side

Same subtasks, same work, same budget. The split buys a smaller largest-context and pays for it twice - in total tokens, and in subtasks that needed something sitting in another agent's window. Push the coupling up and watch which column stops being acceptable first.

One agent

Everything in one context. Nothing is ever invisible to anything else.

Largest context11,200 / 64,000
Answered correctlyNeeded something it could not see
Topology
one agent
Subtasks
8 / 24
Largest context
11,200 tok
Total billed
11,200 tok
Correct
100.0%

Parent + 4 subagents

Five contexts, each paying the briefing again, each blind to the others.

Largest context5,800 / 64,000
Answered correctlyNeeded something it could not see
Topology
parent + 4
Subtasks
8 / 24
Largest context
5,800 tok
Total billed
27,200 tok
Correct
87.5%
Lost to isolation
1
Measured at tick 8 - both sides, same seed, same inputs
MeasureOne agentParent + 4 subagentsGap
Largest single context1120058001.9×
Total tokens billed11200272002.4×
Subtasks answered correctly1.000.881.1×
tick 8
Break it

System prompt, tool definitions, task description. Every context pays it again.

Share that need something another subtask found out.

The decision, as a procedure

Three steps, in order, and the order matters.

Compute the single-agent peak. Briefing plus subtasks times work per subtask. If it fits your window comfortably, stop - you have no problem that subagents solve, and splitting will cost you tokens and correctness for nothing.

If it does not fit, estimate the coupling. How many of these subtasks need what another one found? Independent work - search this list of files, summarise each of these documents - splits beautifully. Work where step nine depends on what step four turned up does not, and no amount of parent-level orchestration fixes it, because the parent only ever sees summaries.

Then pick the smallest fan-out that fits. Every extra agent buys less headroom than the last and costs another full briefing, another report back, and a larger share of coupled subtasks landing in the wrong window.[3]paperReAct: Synergizing Reasoning and Acting in Language ModelsYao, S. et al., ICLR 2023, 2023

The dial

You gainEach context stays small, so tasks that could never fit in one window become possible
You payEvery context pays the briefing again, every subagent pays to report back, and none of them can see what the others learned

The line - when someone asks

Splitting a task across subagents does not make it cheaper, it makes each context smaller - and those are very different things. Every child pays the full briefing again and pays again to report back, so four subagents cost roughly 70% more tokens than doing the same work in one context while cutting the largest single context to about a third. The part that is easy to miss is correctness: a subtask that needs what another subtask discovered will find it sitting in a window it cannot see, and that failure is silent, because the subagent answers confidently from what it does have.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary

Secondary