The MCP Handbook

Chapter 8: The Tool Trap - Problems with Too Many Tools

In the previous chapters we have learned how easy it is to create MCP servers and tools. But beware: more is not always better. In this chapter we discuss the phenomenon of "tool fatigue" in LLMs and how to deal with it.

The Problem: When the Model Can't See the Forest for the Trees

Imagine you connect 20 MCP servers, each with 50 tools. The LLM now has 1,000 tools to choose from. This leads to three critical problems:

  1. Context Window Inflation: Every tool definition (name, description, parameters) consumes tokens. With 1,000 tools, a large part of the model's memory (context window) is already occupied before the user has even posed the first question.
  2. Model Confusion (Hallucinations): The more tools with similar names or descriptions exist, the more likely the model is to pick the wrong tool. It starts to "guess" or mixes up parameters.
  3. Latency: The model needs more time to process the huge list of tools before it can start generating. Response times increase noticeably.

Strategies for a Solution

How do we handle this complexity? There are three proven approaches in the MCP ecosystem:

1. Relevance Filtering (Client-Side)

An intelligent MCP client does not inject all tools at once. Instead, it uses a smaller, faster model (or a vector search) to select only the 5-10 most relevant tools based on the user request and passes them to the main LLM.

2. Hierarchical Tools (Routing)

Instead of offering 100 specialised tools, the server offers a single "router tool".

  • Bad: get_sales_2023, get_sales_2024, get_sales_prognosis...
  • Good: One tool query_data(category, year) which internally decides on the server which logic to call.

3. Clear Delimitation via Prompts

As learned in Chapter 6, Prompts can help the model focus. A Prompt can tell the model: "Today you are only responsible for accounting. Ignore all tools that have nothing to do with finance."

4. Slash-Commands & Deliberate Tool Selection

A very effective pattern in modern chat clients is the combination of user interaction and technical filtering.

Instead of requiring the LLM to guess the right tool out of a huge list, the user sets the direction via a slash command (e.g. /execute_js). This enables a highly optimised flow:

  1. Precise Identification: The moment the user types /execute_js, the client knows exactly which tool will be needed in the next step.
  2. Context Enrichment via prompts/get: The client fetches the matching Prompt for this tool from the server. This often contains specific system instructions or helper text.
  3. Reduced Focus: The client now sends to the LLM only:
    • A small list of standard tools (baseTools).
    • The tool chosen by the user as the forcedTool (the model is instructed to use exactly this tool).
    • The text from the prompt fetch as an additional instruction.

Advantage: The token load drops dramatically, latency is minimised, and the probability that the model picks a wrong tool goes towards zero.

5. Agent Skills & Progressive Disclosure

A modern approach, used both in advanced agent systems (such as Claude Code, Antigravity, or Gemini CLI) and natively over the MCP wire (via the io.modelcontextprotocol/skills extension / SEP-2640), is the Agent Skills pattern.

Here, expert instructions and specialised tools are not permanently loaded into the system prompt. Instead, the principle of Progressive Disclosure (staged revelation) is used:

  1. Discovery (Level 1): Only the name and short description of the available skills sit in the system prompt or arrive via skills/list (low token cost, ~100 tokens per skill).
  2. Activation (Level 2 & 3): When the model recognizes that it needs a skill, it fetches the full SKILL.md (locally or via resources/read on skill://...) and receives complete expert context at exactly the right moment.

This approach keeps the context clean and only focuses the model on details when they are truly relevant. (Details in Chapter 23: Agent Skills.)

Alongside the token load, every multiplication of tools also has a security side: more tools the model can call means more attack surface. How to correctly authenticate and authorise MCP servers is covered in Chapter 17: Security & Authentication.


Testing Tool Overload with `mcp-tester`

With our tool you can check exactly how your server behaves when it is under load or provides a very large number of definitions.

Tool Discovery Benchmarking

Check how long the handshake takes when your server returns many tools:

time ./bin/mcp-tester list --profile heavy-server

Ambiguity Test

Create a test script that tries to confuse the model. Does your server offer two very similar tools? Use mcp-tester to verify that the model (or your script flow) reliably picks the right tool.

The Caching Problem with Dynamic Tools

Many MCP servers offer dynamic tools that can change depending on the context. This can cause problems when the client caches tools. In particular, providers take different approaches to how they handle dynamic tools, and how they implement effective caching. There is the problem of (chat) prefix caching: in essence you send the complete conversation history each time. As long as the tools do not change, that is no problem. But when the tools change, the conversation history must be adjusted - which means that an existing cache can no longer be reused.

The art lies in designing the tools so that they rarely change, and in shaping the conversation flow so that the cache can be reused. Optimised chat workflows therefore make sure, for example, that tool lists stay stable.

There are also "proprietary" approaches to work around this problem. For example, some providers expose cache hints that let you tell them where stable boundaries lie.

Conclusion

A good MCP server is not distinguished by the number of tools, but by their quality and unambiguity. Every tool should fulfil a clear, unique purpose. If you notice that your model makes mistakes, it is often time to slim down the tool list or to formulate the descriptions more precisely.

← Chapter 7: The Power of Combination | Table of Contents | Next Chapter: Return Values →


Copyright Michael Lechner - 2026-04-26

Licence: CC BY-NC 4.0