Chapter 20: Agentic Servers & Sampling - When the Server Asks the AI
So far we have met MCP in this form: the client (the AI) asks the server for tools or data. An agentic server turns this around: it uses a language model itself to solve its task – it plans, calls its own tools, and evaluates intermediate results. To the calling agent, it thereby becomes a sub-agent.
MCP offers two ways to do this:
- Sampling – the server borrows the client's model (
sampling/createMessage). - Own model – the server talks to an LLM directly, for example a local Qwen via Ollama.
As of specification
2026-07-28: Sampling is deprecated (SEP-2577). It remains in the specification for at least twelve months, but new implementations should no longer use it and should integrate LLM provider APIs directly instead. In addition, server-to-client requests such assampling/createMessageare now delivered only via Multi Round-Trip Requests (MRTR, SEP-2322) – the former path, in which the server itself sent a request to the client, has been removed. This chapter therefore describes Sampling for existing implementations and shows the recommended approach with an own model in Section 5.
How solid is the example? The code-review guard in Section 5 is a working teaching example, not a finished product:
- Runs and is tested: The Go code compiles against
modelcontextprotocol/go-sdkv1.8.0. We ran the agent loop in October 2026 against Ollama withqwen3-coder:30band two other Qwen models – the duplicate in the test case was detected in every run (details in Section 5.4). The task variant (Section 5.5) runs too: with themcptaskspackage from the mcp-tester, checked withmcp-tester tasks(all checks passed) and end to end withmcp-tester call --taskagainst the real model.- Only sketched: The similarity index is an interface (
CodeIndex); in the test, a mini index simply returned the files of one directory as candidates. A real search across a repository – for example with code embeddings and a vector database – is yours to add. Configuration (model and URL are constants) and unit tests are missing as well.- Not shown: Deciding per request (few files synchronously, many as a task) and custom progress texts – both need your own task store as in Chapter 19, because
mcptasksdecides per tool and sets thestatusMessageitself. The sampling examples in Section 2 are JSON from the specification, not tested code.- Simplified: The line counter only knows Go comments and does not handle code and a comment on the same line separately.
1. What Is Sampling?
Sampling allows an MCP server to take a "sample" of the model's intelligence. The server sends a list of messages and instructions to the client, and the client returns a response generated by the LLM.
The original arguments in its favor:
- No API keys of its own: The server uses the AI connection the client already has.
- Control stays with the client: The client can block or filter sampling requests, or present them to the user for approval.
- Model choice stays with the user: The client decides which model answers – including a local one.
2. The `sampling/createMessage` Method
Since 2026-07-28, the server answers a client request (tools/call, prompts/get or resources/read) with an InputRequiredResult that contains the sampling request. The client lets the model respond and retries the original request with the answer in inputResponses. If the tool call runs as a task (Chapter 19), the sampling request arrives via inputRequests in tasks/get instead, and the answer is sent via tasks/update.
{
"jsonrpc": "2.0",
"id": 1,
"result": {
"resultType": "input_required",
"inputRequests": {
"summary": {
"method": "sampling/createMessage",
"params": {
"messages": [{ "role": "user", "content": { "type": "text", "text": "Summarize this data: ..." } }],
"systemPrompt": "Be precise and use Markdown tables.",
"modelPreferences": { "intelligencePriority": 0.8, "speedPriority": 0.3 },
"maxTokens": 800
}
}
},
"requestState": "AEAD-protected state"
}
}The most important parameters:
messages: The conversation history (rolesuser,assistant).systemPrompt: Instructions for this generation. The client may modify or ignore it.modelPreferences: What matters to the server –intelligencePriority,speedPriority,costPriority(each 0–1) andhintswith model names. The client makes the choice.maxTokens: Required field; the client must respect it.tools/toolChoice: The server can offer its own tools to the model. If the model responds withstopReason: "toolUse", the server executes the tools and sends a new sampling request with thetool_resultblocks – an agent loop running on the client's model. This requires thesampling.toolscapability.
The server stores its intermediate state in requestState, which the client sends back unchanged. Because this value passes through the client, the server must treat it like input from an attacker and protect it against tampering with HMAC or AEAD.
3. The "Human-in-the-Loop"
The specification places great emphasis on sampling not happening unnoticed in the background. A good client should:
- Inform the user: "Server X wants to use the model to generate a message."
- Allow the user to review and edit the system prompt and messages.
- Present the response for review before passing it on to the server.
- Leave cost control with the user.
This is exactly where the weakness for agentic workflows lies: an agent loop with ten rounds means up to ten approvals. That is safe, but hardly practical for a sub-agent that is supposed to check things quietly in the background – one of the reasons why the direct path to the model is recommended today.
4. Sampling or Own Model?
| Criterion | Sampling (client's model) | Own model in the server |
|---|---|---|
Status in 2026-07-28 |
deprecated | recommended |
| API keys / model operation | none in the server | server needs access (cloud key or local model) |
| Who chooses the model? | client or user | server operator |
| Costs | borne by the client's quota | borne by the server operator (local: hardware only) |
| Approvals | per request by the human | once: trust in the server |
| Agent loop | via MRTR rounds, each round a client pass | directly in the server, no detour |
| Reproducibility | model may differ per client | fixed model, fixed prompts |
The former showcase example for sampling – processing sensitive data with a local model so that it never leaves the network – can actually be implemented more directly with an own model: the server talks to the local model itself, and the data never leaves the machine in the first place.
5. The Server as a Sub-Agent: a Code-Review Guard with a Local LLM
5.1 The Scenario
Claude (or another coding agent) writes code. A second, specialized agent is supposed to check every change:
- a) Does this code already exist? Agents readily write a third variant of
parseConfiginstead of finding the existing one. - b) Are the in-house standards being followed? For example: at most 250 lines of code per file, comments don't count.
The reviewer should use a local model (such as Qwen via Ollama): the code never leaves the machine, each check costs only electricity, and the main agent's model doesn't have to fill its context with half the repository.
Built as an MCP server, this reviewer is a sub-agent that any MCP-capable host can use – Claude Code, Cursor, opencode or your own application.
5.2 The Architecture
Main agent (Claude) MCP server "code-review"
─────────────────── ────────────────────────────
writes handler.go
tools/call review_code ──────────► ① rules in code
{files: ["handler.go"]} code lines ≤ 250?
◄─ CreateTaskResult (working) ──── ② agent loop with Qwen
searches similar code
tasks/get ───────────────────────► (search_similar_code
◄─ working, statusMessage ──────── → index)
judges candidates
tasks/get ───────────────────────► ③ findings as JSON
◄─ completed, {findings: […]} ────
fixes the findingsTwo design decisions carry the whole thing:
1. Compute, don't think. The server checks the 250-line rule with code, not with the LLM. Counting is deterministic, fast and error-free; a language model would miscount. This is the pattern from Chapter 4: the model only gets what requires judgment.
| Check | Who checks | Why |
|---|---|---|
| Lines of code per file | Go code | exactly countable |
| Naming, formatting | linter (golangci-lint) |
rule set already exists |
| Finding duplicate candidates | index (embeddings, AST hashes) | search across the whole repository |
| "Is this really the same logic?" | local LLM | requires understanding |
| Soft rules ("one responsibility per file") | local LLM | requires understanding |
2. Separate searching from judging. The model does not search the repository itself. An index supplies candidates – for example via a local embedding model and a vector database, or via an existing code-graph server. The model only decides whether a candidate is truly equivalent. This keeps the context small, which matters especially for local models with a limited context window.
5.3 Implementation in Go
The server follows the in-house standard: Go with modelcontextprotocol/go-sdk. Ollama provides an OpenAI-compatible interface through which the model can call tools. The example is deliberately compact: the line counter only knows Go comments, and CodeIndex stands for any similarity index.
package main
import (
"bufio"
"bytes"
"context"
"encoding/json"
"fmt"
"net/http"
"os"
"strings"
"github.com/modelcontextprotocol/go-sdk/mcp"
)
const (
maxCodeLines = 250 // in-house standard: lines of code per file, comments don't count
ollamaURL = "http://localhost:11434/v1/chat/completions"
reviewModel = "qwen3-coder:30b"
maxTurns = 6 // upper limit for the agent loop
)
type ReviewIn struct {
Files []string `json:"files" jsonschema:"paths of the changed files"`
}
type Finding struct {
File string `json:"file"`
Rule string `json:"rule"` // e.g. "max-lines", "duplicate", "naming"
Severity string `json:"severity"` // "error" | "warning"
Message string `json:"message"`
}
type ReviewOut struct {
Findings []Finding `json:"findings"`
}
// CodeIndex finds similar code in the repository (embeddings, AST hashes, ...).
type CodeIndex interface {
Similar(ctx context.Context, code string, limit int) ([]string, error)
}
// countCodeLines counts lines of code excluding blank lines and comments – deterministic, no LLM.
func countCodeLines(src string) int {
n, inBlock := 0, false
sc := bufio.NewScanner(strings.NewReader(src))
for sc.Scan() {
l := strings.TrimSpace(sc.Text())
switch {
case inBlock:
inBlock = !strings.Contains(l, "*/")
case strings.HasPrefix(l, "/*"):
inBlock = !strings.Contains(l, "*/")
case l == "", strings.HasPrefix(l, "//"):
default:
n++
}
}
return n
}
// Stage 1: check hard rules in code ("compute, don't think").
func checkRules(path, src string) []Finding {
if n := countCodeLines(src); n > maxCodeLines {
return []Finding{{File: path, Rule: "max-lines", Severity: "error",
Message: fmt.Sprintf("%d lines of code (allowed: %d) – split along a seam", n, maxCodeLines)}}
}
return nil
}
type chatMsg struct {
Role string `json:"role"`
Content string `json:"content"`
ToolCalls []toolCall `json:"tool_calls,omitempty"`
ToolCallID string `json:"tool_call_id,omitempty"`
}
type toolCall struct {
ID string `json:"id"`
Type string `json:"type"`
Function struct {
Name string `json:"name"`
Arguments string `json:"arguments"`
} `json:"function"`
}
// The sub-agent's only tool: search the repository for similar code.
var agentTools = []map[string]any{{
"type": "function",
"function": map[string]any{
"name": "search_similar_code",
"description": "Searches the repository for code similar to the given snippet.",
"parameters": map[string]any{
"type": "object",
"properties": map[string]any{"code": map[string]any{"type": "string"}},
"required": []string{"code"},
},
},
}}
func chat(ctx context.Context, msgs []chatMsg) (chatMsg, error) {
body, _ := json.Marshal(map[string]any{"model": reviewModel, "messages": msgs, "tools": agentTools})
req, err := http.NewRequestWithContext(ctx, http.MethodPost, ollamaURL, bytes.NewReader(body))
if err != nil {
return chatMsg{}, err
}
req.Header.Set("Content-Type", "application/json")
resp, err := http.DefaultClient.Do(req)
if err != nil {
return chatMsg{}, err
}
defer resp.Body.Close()
var out struct {
Choices []struct{ Message chatMsg } `json:"choices"`
}
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil || len(out.Choices) == 0 {
return chatMsg{}, fmt.Errorf("invalid response from model: %v", err)
}
return out.Choices[0].Message, nil
}
// Stage 2: agent loop with the local model – judge duplicates and soft rules.
func reviewWithLLM(ctx context.Context, idx CodeIndex, path, src string) ([]Finding, error) {
msgs := []chatMsg{
{Role: "system", Content: "You review Go code for duplicates and violations of the in-house rules. " +
"Use search_similar_code for every new function. Allowed rules: duplicate, responsibility, naming. " +
`At the end, answer ONLY with JSON: {"findings":[{"rule":"...","severity":"error|warning","message":"..."}]}`},
{Role: "user", Content: "File " + path + ":\n\n" + src},
}
for turn := 0; turn < maxTurns; turn++ {
msg, err := chat(ctx, msgs)
if err != nil {
return nil, err
}
msgs = append(msgs, msg)
if strings.Contains(msg.Content, "<function=") { // tool call as text instead of tool_calls
msgs = append(msgs, chatMsg{Role: "user", Content: "Call tools via the tool interface, not as text."})
continue
}
if len(msg.ToolCalls) == 0 { // done: the model delivers its verdict
var out ReviewOut
if err := json.Unmarshal([]byte(extractJSON(msg.Content)), &out); err != nil {
return nil, fmt.Errorf("verdict is not JSON: %w", err)
}
for i := range out.Findings {
out.Findings[i].File = path
}
return out.Findings, nil
}
for _, tc := range msg.ToolCalls { // execute tools, return results
var args struct{ Code string }
_ = json.Unmarshal([]byte(tc.Function.Arguments), &args)
hits, err := idx.Similar(ctx, args.Code, 5)
result := strings.Join(hits, "\n---\n")
if err != nil {
result = "Error: " + err.Error()
}
msgs = append(msgs, chatMsg{Role: "tool", ToolCallID: tc.ID, Content: result})
}
}
return nil, fmt.Errorf("no verdict after %d rounds", maxTurns)
}
// extractJSON cuts the JSON object out of the response – local models like to wrap it in ```json fences.
func extractJSON(s string) string {
if i, j := strings.Index(s, "{"), strings.LastIndex(s, "}"); i >= 0 && j > i {
return s[i : j+1]
}
return s
}
func register(s *mcp.Server, idx CodeIndex) {
mcp.AddTool(s, &mcp.Tool{
Name: "review_code",
Description: "Checks changed files against the in-house standards (max. 250 lines of code per file) " +
"and searches the repository for existing, equivalent code. " +
"Call after every change to Go files and fix the findings.",
}, func(ctx context.Context, _ *mcp.CallToolRequest, in ReviewIn) (*mcp.CallToolResult, ReviewOut, error) {
var out ReviewOut
for _, path := range in.Files {
src, err := os.ReadFile(path)
if err != nil {
return nil, ReviewOut{}, err
}
out.Findings = append(out.Findings, checkRules(path, string(src))...)
llm, err := reviewWithLLM(ctx, idx, path, string(src))
if err != nil {
out.Findings = append(out.Findings, Finding{File: path, Rule: "review", Severity: "warning",
Message: "LLM review not possible: " + err.Error()})
continue
}
out.Findings = append(out.Findings, llm...)
}
return nil, out, nil // findings are a result, not an error
})
}Three details matter more than they appear:
- The loop is bounded (
maxTurns). A local model that searches in circles must not keep the main agent waiting forever. - If the model fails, the server still delivers a result. The deterministic rules always apply; the LLM part is reported as a warning rather than an error.
- Findings are not an error. The tool returns
isError: false– even with ten violations. The check worked; its result is the findings. AsstructuredContent(viaReviewOut), the main agent can work through them one by one.
5.4 Lessons from the Test Run
We ran the loop against several local Qwen models under Ollama. Test case: a new function parseSettings that is almost line-for-line identical to an existing LoadConfig; the index returns LoadConfig as a candidate. All models detected the duplicate. Along the way, however, exactly the problems surfaced that you have to plan for with local models – which is why the code above contains two safeguards:
| Observation | Frequency | Countermeasure |
|---|---|---|
qwen3-coder:30b writes the tool call in its own format (<function=…>) as text instead of returning tool_calls |
about every third run | correction round: detect the text, send a hint to the model, continue the loop |
The verdict arrives in a Markdown block (```json … ```) instead of as plain JSON |
model-dependent | extractJSON cuts out the object |
| Rule names vary ("Duplicated code", "NoDuplication", "DRY" …) | with every model when not prescribed | fixed list of allowed rule IDs in the system prompt |
Additional, varying warnings (naming, responsibility) for an unremarkable function |
in about half of the runs | treat LLM findings only as warning; error stays reserved for the deterministic rules and clear duplicates |
Runtime: 4–8 s (qwen3-coder:30b, MoE) to 30–60 s (dense 27–35B models) per file |
– | run as a task once there are several files (Section 5.5) |
With both safeguards in place, qwen3-coder:30b ran through flawlessly in all repetitions. The lesson is more general than this example: a local sub-agent needs the same robustness as any other unreliable interface – validate the format, retry a bounded number of times, and design the result so that a slip by the model never spoils the deterministic part.
For the index itself, a local code embedding model (such as jina-embeddings-v2-base-code) together with a vector database works well; alternatively, an existing code-graph MCP server can be connected.
5.5 Synchronous or as a Task?
The Tasks extension lets the server decide per request (Chapter 19). That fits perfectly here:
- One file, small change: respond synchronously. The check takes a few seconds; a task would be overhead.
- Many files or a large repository: return a
CreateTaskResultand move the work into the task store from Chapter 19. ViastatusMessage, the main agent sees what is currently happening ("File 3/12: searching for duplicates …") and can keep working in the meantime. - Need to ask back? If the reviewer finds a function that is 95% identical to an existing one, it can ask via
input_requiredand an elicitation: "Should I suggest the existing function or accept the new one?"
Because the server only returns a CreateTaskResult when the request declares the extension, the same server also works with hosts that don't support tasks yet – they simply get the answer synchronously.
Implementation with `mcptasks`
The go-sdk does not ship a task implementation yet, but it offers the necessary extension points (Chapter 19, "Connecting to the go-sdk"). The mcptasks package from the mcp-tester uses them; with it, the review server becomes task-capable without touching the handler:
func main() {
caps := &mcp.ServerCapabilities{Tools: &mcp.ToolCapabilities{}}
mcptasks.Declare(caps)
s := mcp.NewServer(&mcp.Implementation{Name: "code-review", Version: "0.1.0"},
&mcp.ServerOptions{Capabilities: caps})
register(s, newIndex()) // review_code from Section 5.3, newIndex: your CodeIndex
if err := mcptasks.Enable(s, mcptasks.NewStore(), "review_code"); err != nil {
log.Fatal(err)
}
if err := s.Run(context.Background(), &mcp.StdioTransport{}); err != nil {
log.Fatal(err)
}
}The test run with the mcp-tester (the test ran with the German version of the prompt, so the model answered in German):
$ mcp-tester call review_code --task -c ./review-server --args '{"files":["handler/settings.go"]}'
[TASK] a67c90d1… working: the operation is in progress
[TASK] a67c90d1… completed: the operation completed
StructuredContent:
{ "findings": [
{ "file": "handler/settings.go", "rule": "duplicate", "severity": "error",
"message": "Die Funktion parseSettings … ist eine doppelte Implementation von LoadConfig in config/load.go. …" },
{ "file": "handler/settings.go", "rule": "responsibility", "severity": "warning", "message": "…" } ] }mcp-tester tasks --tool review_code … additionally checked the error codes, the task handle, ID entropy, durable creation, every tasks/get and the status transitions – all checks passed.
Two limits of mcptasks are worth knowing: the package always turns a tool into a task when the client declares the extension – so "few files synchronously" is not possible with it. And it sets the statusMessage itself; "File 3/12 …" needs your own store following the pattern from Chapter 19.
5.6 How the Main Agent Uses the Reviewer
A sub-agent is only useful if it actually gets called. There are three levels for this, from soft to hard:
- Tool description: "Call after every change to Go files" is right there in the
description. The model reads it on every tool selection. - Project instruction: A sentence in
AGENTS.mdorCLAUDE.md("Before every commit, runreview_codeon all changed files and fix allerrorfindings") turns the call into a working rule. - Enforcement via the host: Some hosts offer hooks that automatically run a command after every write operation (such as Claude Code's
PostToolUsehooks). That way, the check no longer depends on the model remembering it. This is, however, host functionality, not MCP.
6. MCP Sub-Agent or Host Sub-Agent?
Many hosts come with their own sub-agents (Claude Code, Agent SDKs, IDE agents). When is the MCP route worth it?
| Criterion | Host sub-agent | Sub-agent as MCP server |
|---|---|---|
| Model | usually the host's model | freely selectable, including local (Qwen, Llama, …) |
| Data | goes to the host's model provider | stays on your own machine if desired |
| Cost per call | tokens of the main model | local: close to zero |
| Reuse | tied to one host | any MCP-capable host |
| Tools | those of the host | its own, narrowly tailored ones (index, linter) |
| Deterministic parts | hard to guarantee | guaranteed in the server code |
| Setup | one configuration file | run a server, provide a model |
| Quality | top-tier model | local model – usually sufficient for narrowly scoped checks |
As a rule of thumb: open-ended, creative subtasks ("research", "design") belong to the host sub-agent with a strong model. Narrowly scoped, recurring checks with fixed rules – standards, duplicates, licenses, security patterns – are ideal MCP sub-agents: they benefit from deterministic code, their own tools and an inexpensive local model, and they behave the same in every host.
Conclusion
Agentic servers turn MCP servers from passive toolboxes into active collaborators. With 2026-07-28, the recommended approach has shifted: away from sampling via the client's model, toward the server that brings its own model – preferably a local one. Combined with the Tasks extension from Chapter 19, this yields a full-fledged sub-agent: the main agent delegates, the server checks with a mix of code and model, and the result comes back structured.
← Chapter 19: Tasks | Table of Contents | Next Chapter: Elicitation →
Copyright Michael Lechner – 2026-10-09 (Sampling deprecated & MRTR per specification 2026-07-28, new section "The Server as a Sub-Agent")