The MCP Handbook

Chapter 14: Quality Assurance – The Server Inspector & Quality Scoring

Writing a basic MCP server can be done in an afternoon. Building a resilient, specification-compliant, and AI-optimized MCP server, however, is true systems engineering. In this chapter, we explore how to use the open-source mcp-tester to objectively evaluate your implementation, detect specification changes early, and systematically optimize server quality.


1. Why Is Quality Mission-Critical for MCP?

Large Language Models do not possess human common sense: they depend entirely on the metadata and schemas advertised by an MCP server during handshake and invocation. Incomplete or outdated responses trigger direct failures:

  • Missing Tool Descriptions: The model cannot determine which tool fits the task, causing it to guess or hallucinate.
  • Missing Output Schemas: When tools return arbitrary prose instead of validated JSON (see Chapter 9), the model must guess the schema on the fly—wasting tokens and introducing runtime parsing bugs.
  • Missing Prompts & Guidelines: The model lacks persona definitions or operational guardrails for chaining your tools (see Chapter 6).

2. The `mcp-tester`: Diagnostic Tool for Developers

To eliminate these blind spots, we developed mcp-tester. It acts as a specialized, protocol-strict MCP client that validates every facet of your server:

MCP Server Quality Inspector & Scoring


3. Key Advantages of `mcp-tester`

Why is testing with Claude Desktop or Cursor insufficient? The mcp-tester offers decisive architectural benefits:

A. Early Detection of Specification Changes (Spec Drift Detection)

The Model Context Protocol evolves continuously:

  • Modern transport standards (Streamable HTTP replacing legacy dual-endpoint SSE),
  • Protocol extensions (such as Tasks via SEP-2663 or Skills over MCP via SEP-2640),
  • Stricter capability declarations and metadata requirements (icons, tool annotations).

When the official specification advances or an underlying SDK ages, standard LLM clients often fail silently or silently drop tools. mcp-tester validates the wire protocol directly against the latest specification. Deprecated JSON-RPC structures, missing headers, or malformed schemas are flagged immediately—long before end users encounter cryptic errors in real conversations.

B. Quality Score (0–100) as an Optimization Roadmap

A binary "Pass/Fail" indicator provides little guidance. A server can be syntactically valid yet poorly optimized for AI interaction. The tester's inspect command computes a Quality Score from 0 to 100 points, generating an actionable diagnostic checklist:

./bin/mcp-tester inspect --profile my-server

Scoring Breakdown:

  1. Missing Prompts (-20 points):
    A production server should register system prompts explaining how its tools are intended to be orchestrated.
  2. Missing Tool Descriptions (-5 points per tool):
    Any tool lacking a clear, descriptive purpose is penalized heavily, as this is the primary cause of model confusion.
  3. Missing Output Schemas (-2 points per tool):
    Structured JSON return schemas are essential for reliable machine parsing.
  4. Visual Metadata (Icons):
    Verifies that referenced icon URIs resolve cleanly and use valid MIME types.

If a server scores 73/100, the inspector outputs the exact reasons:

"-20 pts: No prompts registered; -5 pts: Tool 'export_db' lacks a description; -2 pts: Tool 'query' missing structured output schema."
This gives developers an immediate, prioritized punch list.

C. Black-Box Integration Testing vs. Unit Testing

A unit test in Go or Python tests internal function logic (handleAdd(1, 2) == 3).
In contrast, mcp-tester exercises end-to-end protocol behavior from the outside:

  • Handshake negotiation over Stdio pipes or Streamable HTTP,
  • JSON-RPC serialization and schema validation,
  • Timeout handling and cancellation propagation.

4. Validating Visual Metadata (Icons)

Modern agent interfaces and IDEs display visual icons for tools, resources, and prompts. The tester provides dedicated inspection flags to verify asset links:

# Verify reachability of all icon URIs
./bin/mcp-tester list --check-icons --profile my-server

# Download all icons to a local directory for visual inspection
./bin/mcp-tester list --download-icons "./icons_export" --profile my-server

5. Checklist for a "Perfect Server" (100/100 Score)

To achieve a full 100/100 score in mcp-tester inspect, ensure your server satisfies:

  • Handshake & Version: Negotiates the latest protocol revision and advertises capabilities accurately.
  • Tool Descriptions: Every tool provides detailed purpose, constraints, and trigger keywords.
  • Strict Typing: Input parameters and output payloads are defined using standard JSON schemas.
  • Prompts Defined: At least one prompt sets the system context or defines a workflow for the model.
  • Icon Integrity: All declared icon URLs resolve successfully.

6. The Ideal Balance: Unit Tests & MCP Tester

In production, combine both approaches into a coherent testing pipeline:

┌────────────────────────────────────────────────────────┐
│                      Unit Tests                        │
│   Validate internal logic & edge cases in milliseconds │
└───────────────────────────┬────────────────────────────┘
                            │
                            ▼
┌────────────────────────────────────────────────────────┐
│                   mcp-tester inspect                   │
│   Validate protocol compliance, schemas & score        │
└───────────────────────────┬────────────────────────────┘
                            │
                            ▼
┌────────────────────────────────────────────────────────┐
│              mcp-tester test (CI/CD scripts)           │
│   Automate regression testing before every deployment  │
└────────────────────────────────────────────────────────┘

Example: Unit Test for a Go Tool Handler

func TestHandleAdd(t *testing.T) {
    ctx := context.Background()
    args := map[string]any{"a": 10.0, "b": 20.0}

    result, structured, err := handleAdd(ctx, nil, args)
    if err != nil {
        t.Fatalf("Unexpected error: %v", err)
    }

    if structured.(map[string]any)["sum"] != 30.0 {
        t.Errorf("Unexpected result: %v", structured)
    }
}

Rule of Thumb: Use unit tests for internal calculation logic and mcp-tester for protocol compliance, schema fidelity, and agent interaction quality.


Conclusion

Quality assurance in MCP means eliminating friction between servers and AI agents. By utilizing mcp-tester (github.com/hmsoft0815/mlc_mcptester), you detect specification drift early and receive a transparent scoring roadmap to elevate your server to enterprise standards.

← Chapter 13: The Artifact Pattern | Table of Contents | Next Chapter: Automation & CI/CD →


Copyright Michael Lechner – 2026-09-27 (Updated for MCP Quality Scoring, Spec Drift Detection & GitHub Open Source)

Licence: CC BY-NC 4.0