Giving one AI agent a massive tool list crashes its accuracy to roughly one pick in seven. In a recent test, a standard agent handed a giant catalogue chose the right tool just 13.62% of the time [1]. Filtering that list first lifted accuracy to 43.13% and cut prompt tokens in half, simply by showing the model fewer choices [1].
The big picture:
Tool choice is just a pick from a list, and models break down as that list grows. Recent studies hit the same wall again and again. They prove that tool count is a strict limit, not a feature list.
The scale of these lists is now huge. A recent scan of the official Model Context Protocol (MCP) registry found 98,291 tools across its servers [4]. MCP acts as the standard bridge connecting agents to company systems, opening the door to a massive pool of choices.
The hard limits sit far above the point where quality actually drops. One major vendor caps a single request at 128 tools [5]. A large language model (LLM) will accept that entire list and still make terrible choices.
My read: the ceiling that matters is a clear design choice, not the limit in the vendor’s software.
Why it matters: An agent that picks the wrong tool to change or save data does not fail loudly. It calls a valid endpoint, returns a believable result, and moves on, burying the mistake.
By the numbers
- 13.62% — Baseline accuracy: A standard agent handed a massive tool set chose the correct tool just 13.62% of the time, while a filtered search lifted that to 43.13% [1].
- 50% — Prompt token drop: The same search approach cut prompt tokens by over 50% during an MCP stress test, making the short list cheaper to run [1].
- 7 tools — Adaptive depth: A dynamic search method matched the coverage of showing 50 tools (90.3% versus 90.8%) while presenting just 7 on average [2].
- 98,291 — Registry catalogue: A complete scan of the official MCP registry counted 98,291 tools exposed by its servers [4].
What I’d watch:
What strikes me here is that every proven fix controls what the model sees, rather than waiting for a smarter model. The research groups these controls into three approaches.
- Retrieval-scoped lists: Developers keep an index of tool descriptions, score them against the user request, and pass only the top matches to the model. This method drove the accuracy jump from 13.62% to 43.13% [1].
- Adaptive shortlist depth: Systems let the search dig deeper only when a short list fails. On a 3,251-tool registry, a fixed list of five tools found nothing for hard queries, while a deeper search still found 16.7% [2].
- Smaller specialist agents: Builders split one overloaded agent into narrow workers with tiny tool sets. Internal field notes show a 42-tool agent completed 37% of tasks, while three specialists holding six tools each reached 94.8% completion [6].
I’ve been watching how teams choose between building a gateway to manage the catalogue and redrawing the org chart into specialist agents. I’d watch which path the big vendors ship as their default pattern next year.
The catch
None of this is free, and the numbers carry sharp edges. A short list can easily leave out the right tool entirely. On the hardest queries in a recent depth study, a fixed list of five tools found absolutely nothing [2].
Search requires a dedicated index that teams must build, update, and tune. The measured rates also come from lab tests rather than live traffic. I read the 13.62% figure as a floor showing how bad a massive catalogue can get, not a guaranteed failure rate for every setup.
The part I keep circling is that the core pattern holds even as exact figures shift. A model choosing from many tools performs worse than one choosing from a few. I could be wrong about exactly where the ceiling sits for your specific registry, but handing an agent a hundred tools and hoping it sorts them out is a losing bet.
At a glance
- The Big Shift: An agent’s tool count is a strict design budget, as selection reliability crashes when the list grows.
- Why It Matters: An agent that picks the wrong tool to change or save data fails quietly, hiding the mistake until later.
- What I’d Watch: How engineering teams bound their tool lists. The evidence points to three distinct methods:
- Retrieval-scoped lists: An index scores tools against each request and passes only the top matches to the model.
- Adaptive shortlist depth: The search system digs deeper only when an initial short list misses the right tool.
- Specialist agents: Builders split one overloaded agent into several narrow ones, each holding a tiny tool set.
- The Catch: A short list can leave out the right tool entirely, and building a search index adds heavy overhead. The failure rates also reflect lab tests rather than live traffic.
Related reading
- OpenAI just made the agent loop a commodity
- MCP Skills Extension — more on Autonomous & Agentic Workflows
- Planner-as-Router: Fold the Model Choice Into the Plan — more on Autonomous & Agentic Workflows
Sources
[1] Tiantian Gan and Qiyao Sun, “RAG-MCP: Mitigating Prompt Bloat in LLM Tool Selection via Retrieval-Augmented Generation” (arXiv:2505.03275, submitted 2025-05-06) — the MCP stress test: tool-selection accuracy 43.13% vs a 13.62% baseline, prompt tokens cut by over 50% (https://arxiv.org/abs/2505.03275) [2] Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh and Joey Blackwell II, “How Many Tools Should an LLM Agent See? A Chance-Corrected Answer” (arXiv:2605.24660, v2 2026-06-07) — BFCL (370 tools) learned depth matches showing 50 tools (90.3% vs 90.8%) at 7 shown on average; ToolBench (3,251 tools) a fixed 5 gets 64.7% vs 61.9% coverage but nothing on hard queries, against 16.7% for the deeper search; 93.1% vs 87.1% with Claude Sonnet 4.6 (https://arxiv.org/abs/2605.24660) [3] Dennis Thompson, “When too many tools become too much context”, WRITER engineering — the RAG-MCP team’s own writeup: retrieval “more than triples tool-selection accuracy and reduce prompt tokens by over 50%”, inside an enterprise MCP Gateway that governs and scales the tool catalogue (https://writer.com/engineering/rag-mcp) [4] Artem Trofimov and Boris Novikov, “When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary” (arXiv:2609.15397, submitted 2026-09-14) — the official MCP registry snapshot measured 98,291 tools exposed by registered servers (https://arxiv.org/abs/2609.15397) [5] OpenAI Developer Community, “Maximum amount of tools for the bot to use?” — the documented vendor cap: “an array of a maximum of 128 functions” per request (https://community.openai.com/t/maximum-amount-of-tools-for-the-bot-to-use/665720) [6] Field notes — context/growth_os/founder-voice.md §multi_agent_enterprise_fabric (“Bound autonomous agents to lean, decoupled tool sets (≤ 8 per agent)”) and context/growth_os/customer-truth.md Anecdote 2 (a dispatch agent with 42 tools: 63% parameter hallucination, 37% task completion; three specialists at ≤ 6 tools each: 94.8%)