yield point
Agentsintermediateupdated 2026-08-24

Tool selection at scale

An agent does not fail to pick the right tool because it has too many tools. It fails because three of them look alike, and three lookalikes cost exactly what sixty unrelated tools would.

Everyone building agents hits this and almost nobody has a number for it. You mount another MCP server, and things get worse in a way that is hard to attribute: the agent starts reaching for the wrong thing, occasionally, on tasks it used to get right.

The usual conclusion is “too many tools”. That is close enough to be actively misleading.

Thirty tools and the agent picks right two times in three. Now add three that do almost the same job as the right one, and watch that cost you exactly what mounting sixty more unrelated tools would have.

Right tool chosen
-
Mounted
30 tools
90% budget
7 tools
A duplicate costs
20 tools
Tool definitions
3,600 tok
Wrong calls
0
tick 0 / 3000
Break it
Do

Every tool on every server the agent can see, not just the ones this task needs.

Tools that do almost the same job as the right one. These are the expensive ones.

How much more the right tool looks like the answer than an unrelated one does.

Safety properties
  • ✓
    The right tool is always mounted

    held on every tick so far

  • ✓
    Selection probabilities sum to one

    held on every tick so far

Where the number comes from

Each cell is a mounted tool, shaded by how likely the agent is to pick it. The bright one is correct. Nothing else looks like a threat, and collectively they take a third of the probability.

The rule is Luce’s choice model - what you get when every option carries an independent noisy score and the highest wins, which is also what a softmax computes:

p(right tool) = 1 / (1 + sum over the others of exp(-margin))

The margin is how much more the right tool looks like the answer than a given other one does. That single term is doing all the work, and it is inside an exponential.

It is not the count

Set tools to 30 and watch: about two picks in three. Now mount another server, taking it to 60, none of them related to this task. Accuracy falls to around one in two - so count does cost you something, and irrelevant is not the same as free.

But now put tools back to 30 and press Add a near-duplicate three times.

Change Accuracy
30 distinct tools 65%
90 distinct tools 38%
30 tools, 3 of them near-duplicates 39%

This is the reframe worth taking away. The question is not how many tools do I have. It is how many of them look like each other - and a team that would refuse to mount sixty more tools will ship search, search_v2 and a legacy one nobody deleted over a single sprint, without anyone raising it.

Identical tool counts, identical token cost, identical task. One agent's tools all do different things; four of the other's do nearly the same job. Raise the count on both and watch the gap grow rather than close.

All distinct

Every tool is clearly for something else, so each one competes for almost none of the probability.

Right tool chosen
68%
Mounted
30 tools
90% budget
7 tools
A duplicate costs
20 tools
Tool definitions
3,600 tok
Wrong calls
61

Four overlap

Four tools do nearly the same job as the right one, and each competes with it almost evenly.

Right tool chosen
32%
Mounted
30 tools
90% budget
1 tools
A duplicate costs
20 tools
Tool definitions
3,600 tok
Wrong calls
131
Measured at tick 200 - both sides, same seed, same inputs
MeasureAll distinctFour overlapGap
Right tool chosen0.690.342.0×
Tool definitions36003600same
Wrong tool called61.0131.02.1×
tick 200
Break it
Do

Every tool on every server the agent can see, not just the ones this task needs.

How much more the right tool looks like the answer than an unrelated one does.

The lever is descriptions, not deletion

Deleting tools is the obvious fix and usually the impossible one, because all three are in use somewhere.

Raise Distinctiveness from duplicates instead. Nothing is unmounted, and the accuracy comes back - because you have moved each duplicate from competing almost evenly to competing barely at all. Since the margin sits inside an exponential, small improvements in clarity buy large improvements in selection.

In practice that means:

  • Say when not to use it. A description that only lists capability gives the model nothing to discriminate on. The distinguishing case is the useful sentence.
  • Name the case, not the capability. search_customers_by_email rather than search.
  • Treat two confusable descriptions as a defect. If a careful human reader would have to think about which of two tools applies, the model is doing worse than that reader.

This is unglamorous work and it outperforms nearly anything else you could do with the same afternoon. The literature on scaling tool use to thousands of APIs has landed in the same place from the other direction: retrieval over documented APIs, and documentation quality, rather than simply presenting the model with more options.[1]paperGorilla: Large Language Model Connected with Massive APIsPatil, S. G. et al., 2023[2]paperToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIsQin, Y. et al., ICLR 2024, 2024

The cost you are not counting

Look at tool definitions in the widget. Thirty tools at 120 tokens each is 3,600 tokens; at 120 tools it is 14,400.

Those tokens are not ordinary tokens. Tool definitions are concatenated at the very front of the prompt, which means they are the most expensive block you can invalidate in your cache, and they are pushing everything you actually retrieved further from the positions the model reads reliably. Mounting a server you use twice a week is paying that on every request.

What to carry away

Count how many of your tools look alike. That number predicts your failures better than the total does.

Descriptions are the cheap lever. The margin sits in an exponential, so clarity pays disproportionately, and it needs no deletion or fine-tuning.

Scope the tool list to the task. Mounted-for-everything is a default, not a decision.

Selection failures are not capability failures. The agent had the tool. Before you reach for a better model, delete six tools and rewrite three descriptions.

The dial

You gainEvery tool you mount is a capability the agent can reach for without you anticipating the need
You payEvery tool also competes for the choice, in proportion to how much it resembles the right one

The line - when someone asks

Tool selection follows a choice rule: the right tool's share of the probability is one over one plus the sum of what every other tool contributes, and each contributes exp(-margin) where the margin is how distinguishable it is. That means the count matters far less than the confusability - at a typical margin, one near-duplicate is worth about twenty clearly distinct tools. The lever is therefore not deleting tools, it is writing descriptions that say when not to use them.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary