The MCP Handbook

Chapter 12: The Limits of AI - Model Constraints and Context Overflow

MCP gives us the power to couple huge amounts of data and powerful tools to an LLM. But at the end of the day, the LLM is still the bottleneck. In this chapter we discuss the physical and cognitive limits of models.

1. The Multimodality Problem (Modality Mismatch)

Not every model can "see" or "hear".

  • Scenario: An MCP server sends an image via ImageContent.
  • Model limit: If you use a purely text-based model (like many smaller open-source models or older GPT versions), it cannot make anything of the image.
  • Consequence: The client must either filter the image out, convert it into a text description, or the request is rejected with an error.

Developer tip: Your server should ideally offer alternative text representations for binary data in case the model is not multimodal.


2. Context Overflow (Context Window Overflow)

Every model has a "context window" (e.g. 128k tokens on GPT-4 or 200k on modern SOTA models). Every piece of information that comes in via MCP takes up space in this window.

The Danger of Large Resources

When a tool or a resource returns a 5 MB log file or a huge JSON object, the following happens:

  1. Token inflation: 5 MB of text corresponds to millions of tokens.
  2. Overflow: The context window fills up instantly.
  3. Memory loss: The model starts to "forget" the beginning of the conversation. This often affects the system instructions (Prompts) from Chapter 6. The model loses its persona or ignores safety rules.

3. Cost and Latency

Do not forget: with cloud models you pay per token.

  • Sending a high-resolution image or a massive database response via MCP can make a single chat request very expensive.
  • In addition, latency (time-to-first-token) rises dramatically when the model first has to process hundreds of thousands of tokens before it can answer.

4. Strategies for Robust MCP Systems

To work around these limits, server developers should use the following techniques:

  • Pagination: Never send all data at once. Use lists with nextCursor.
  • Summarisation: Instead of sending 1,000 lines of log, let the server pre-analyse the logs and send only a summary.
  • Sampling preference: Choose the model to match the server. A server that generates many images mandatorily needs a multimodal frontend.

Validation with `mcp-tester`

Use the mcp-tester to keep an eye on the size of your responses. When saving images in script mode, pay attention to their file size. If a resource query takes several seconds in the tester, it will most likely lead to a timeout or massive latency in real LLM chat.

Conclusion

An MCP developer must always keep the token economy in view. Just because you can expose an entire hard drive as a resource does not mean that the LLM should read it. Less is often more - precision beats volume.

← Chapter 11: Real-Time Feedback | Table of Contents | Next Chapter: The Artifact Pattern →


Copyright Michael Lechner - 2026-02-28

Licence: CC BY-NC 4.0