βοΈ Developer (pro) Β· Guide 2 of 2
Building an MCP Server: The Decisions the Spec Leaves to You
11 min read Β· Last reviewed 23 Sep 2026

The Model Context Protocol tells you how to frame a message, declare a tool and return a result. It does not tell you how many tools to expose, what to call them, how much to return, or what to do when a model calls your delete endpoint with a guessed identifier. Those decisions determine whether your server is useful, and they are all yours.
This is not a protocol tutorial. For the wire format and the SDKs, go to modelcontextprotocol.io and the MCP Python SDK. What follows is the judgement layer on top.
How many tools, and how coarse
The instinct is to map your API one-to-one. Thirty endpoints, thirty tools. It is the wrong default, for a measurable rather than aesthetic reason: every tool definition sits in the model's context on every turn, and every near-identical option is another chance to pick the wrong one.
Anthropic's engineering guidance on tool design is blunt about it β "more tools don't always lead to better outcomes" β and recommends consolidating multi-step operations into single tools: a schedule_event tool instead of separate find_availability and create_event tools, because the agent's natural unit of work is the whole task, not the API call. It prefers search-style tools over list-everything tools for the same reason.
The opposite failure is equally real. One execute tool with a free-form action string returns you to the problem the protocol was meant to solve: no schema to reason against, validation by string parsing, and a typo that becomes a silent no-op.
The pattern that survives contact with users is grouping. The GitHub MCP Server ships 24 toolsets β repos, issues, pull_requests, actions, code_security and so on β defaulting to just context, repos, issues, pull_requests and users, selectable via --toolsets or GITHUB_TOOLSETS, plus --tools for individual tools. Its stated reason is the one above: enabling only the toolsets you need helps the model with tool choice and reduces context size. The Kubernetes MCP Server does the same, defaulting to core and config with helm, tekton and others opt-in, and says it improves tool-selection accuracy.
GitHub MCP Server π MCP serverFree
GitHub's official MCP server: issues, pull requests, code, Actions and security alerts
Build coarse tools around tasks, group them into toolsets, ship a small default, and let power users opt in.
Naming and describing for a model, not a colleague
Tool names in the specification should be 1 to 128 characters, are case-sensitive, and should use only ASCII letters, digits, underscore, hyphen and dot. Uniqueness is scoped to your server β the spec notes that clients aggregating several servers may collide on a name like search and should disambiguate themselves, and that your serverInfo name is not reliable for that. Namespace your own names instead of hoping the client does it well.
Descriptions are prompt engineering, not documentation. Write them as onboarding notes for a competent new engineer who has never seen the system: what the tool does, when to reach for it, when not to, what it costs. Anthropic's guidance reports that "even small refinements to tool descriptions can yield dramatic improvements", and that prefix-versus-suffix namespacing measurably changes accuracy β so the only way to settle either is to test against real tasks.
Use unambiguous parameter names (user_id, not user), and state cross-tool dependencies: if create_deployment needs a project id that only list_projects returns, say so in both descriptions.
Input schemas that prevent bad calls
Your inputSchema is the cheapest guardrail you will build, because it acts before the call happens: the model reads it as part of the tool definition and it constrains what gets generated.
Use enums rather than free strings wherever the value set is closed. Mark required fields required. Give every property a description β models read them. Set formats and ranges that make illegal calls unrepresentable rather than merely detectable. For a tool with no parameters, the spec recommends { "type": "object", "additionalProperties": false }.
The subtler decision is identifiers. A tool that takes an opaque id invites the model to invent one. Where you can, accept the human-meaningful thing β a repository slug, a project name β and resolve it server-side. If you must hand out handles that live across calls, follow the specification's stateful-tools guidance exactly: make the handle opaque and high-entropy, validate the caller's authorisation against it on every call rather than treating possession as authentication, state its lifetime in the creating tool's description, and return a recoverable error when it expires.
Define an outputSchema when your result has a stable shape. If you provide one, the spec requires your structured results to conform to it and clients should validate, which turns silent contract drift into a visible failure.
What to return, and how much
You can return unstructured content blocks, a structuredContent value, resource links, or several together. The decision is about who the result is for.
Return semantically meaningful data rather than raw identifiers. Anthropic's tool-writing guidance found that resolving arbitrary UUIDs into interpretable names materially improved precision: owner: "a7f3c2e1-..." gives the model nothing to reason with, owner: "platform-team" does. Include the technical id too when a follow-up call needs it.
Then decide your size policy, because the protocol will not. MCP pagination is defined only for list operations β tools/list, resources/list, prompts/list and resource templates β using an opaque cursor whose page size the server chooses and clients must not assume. There is no protocol-level pagination for tool results. If your tool can return ten thousand rows you invent the mechanism: a limit with a sane default, a cursor or offset argument of your own, filter parameters, and a response_format style enum letting the caller ask for concise or detailed output.
When you truncate, say so in the payload and tell the model what to do about it. "Showing 50 of 4,312 matches; narrow with the since or author filter" is a recoverable state. Silently returning the first 50 is a correctness bug that surfaces as a confidently wrong answer three turns later.
Errors a model can recover from
The spec draws a line that is easy to get wrong in implementation. Protocol errors β unknown tool, malformed request β are JSON-RPC errors, and clients may pass them to the model, but the spec notes they are "less likely to result in successful recovery". Tool execution errors β API failures, validation failures, business logic β belong in a normal result with isError: true, and clients should hand those to the model so it can self-correct.
A failed call is therefore not an exception to be logged. It is a message to a reader who will act on it. The spec's own example sets the standard: "Invalid departure date: must be in the future. Current date is 08/08/2025." It names the field, the rule and the fact needed to fix it. Apply that everywhere: name the missing field, list the valid enum values, name the tool that mints a fresh handle, say which scope is required rather than returning a bare 403.
Read-only modes and destructive actions
Assume someone points your server at production on day one, and decide in advance what they can do by accident.
The protocol offers tool annotations β readOnlyHint, destructiveHint, idempotentHint, openWorldHint β and their defaults assume the worst: unannotated tools are treated as non-read-only, potentially destructive, non-idempotent and open-world. Set them honestly. But know what they are: the spec requires clients to treat annotations as untrusted unless they come from a trusted server, and the protocol's own commentary is unsentimental about the limits β an untrusted server can lie, annotations are not enforcement, and they do not make a model resist prompt injection. Their main use is driving confirmation prompts.
Real enforcement is a server-side mode, and the patterns in the directory are consistent:
- Supabase MCP has a
read_onlysetting that excludes mutating tools, aproject_refthat scopes the server to one project (which also removes theproject_idargument from tool inputs and drops account-level tools), andfeaturesgroups to narrow the surface further. - MCP Server for MySQL disables all write operations by default, requiring explicit
ALLOW_INSERT_OPERATION,ALLOW_UPDATE_OPERATIONandALLOW_DELETE_OPERATIONopt-ins, with per-database control and masking for sensitive fields. - The GitHub server's read-only mode skips write tools even when they are explicitly requested, and the Kubernetes server supports a
denied_resourceslist for blocking resource kinds such as Secrets.
Supabase MCP π MCP serverFreemium
Manage Supabase projects, databases, migrations and edge functions from your agent
MCP Server for MySQL π MCP serverOpen source
Read-only-by-default MySQL access with opt-in writes, SSH tunnels and PII masking
Notice the shared design: read-only removes the tools rather than refusing them at call time. A tool the model cannot see cannot be talked into use.
stdio versus HTTP, and what each implies for auth
Two transports are standard: stdio, where the client launches your server as a subprocess and exchanges newline-delimited JSON-RPC over its standard streams, and Streamable HTTP, where each message is a POST to a single endpoint. The authorisation consequences are not symmetrical, and the spec states them plainly.
HTTP-based servers should follow the authorisation specification β OAuth 2.1, with the server acting as a resource server that must implement protected resource metadata (RFC 9728), clients sending resource indicators (RFC 8707), and the server validating that every token was issued for it. Servers must not accept or forward tokens issued for anything else: token passthrough is named as an anti-pattern and forbidden, because it breaks audience separation and hands the downstream API a confused deputy.
Stdio servers should not implement that flow; they take credentials from the environment. That is simpler, and it is also why a local server is a soft target, since it runs with the user's privileges. If you offer a local HTTP mode for convenience, the spec's advice is to require an authorisation token or use a Unix domain socket rather than leaving an unauthenticated port on localhost. If you ship both β as the GitHub server does, with a hosted endpoint plus a local container image β document which is supported for whom.
Token scoping and secrets
Ask for the least your tools need and make the rest reachable on demand. The spec's security guidance is explicit that publishing every possible scope, or using omnibus scopes like * or full-access, is a mistake: it inflates the blast radius of a stolen token, makes revocation all-or-nothing, and trains users to click through consent screens. The intended shape is a minimal baseline scope plus step-up authorisation β a privileged call returns 403 with WWW-Authenticate: Bearer error="insufficient_scope", scope="...", naming everything that operation needs in one challenge.
Keep secrets in the environment or a secret store, never in tool arguments or results. If you use the parameter-to-header routing mechanism, the spec warns against marking passwords, keys, tokens or PII that way, since header values are visible to intermediaries. Treat your own logs as an exfiltration surface too.
Versioning without breaking users
MCP versions are dates, not semver: the current revision is 2026-07-28, and the previous line runs back through 2025-11-25. That revision removed the negotiation handshake β every request now declares its protocol version, and a server that does not support it returns an UnsupportedProtocolVersionError listing what it does. Servers can be "dual-era", answering both styles, and the spec publishes a compatibility matrix.
Your own versioning is a separate problem, with tighter constraints than a normal API because your consumer is a model that read your tool list once. The spec requires your tool set not to vary per connection β it may vary by the authorisation presented β and recommends a deterministic ordering so clients can cache the list.
The rules that follow: renaming a tool is breaking, because prompts, evals and user habits reference the name. Removing a parameter is breaking; adding an optional one usually is not. Tightening an enum breaks existing callers. When you must break something, add the new tool alongside the old, mark the old one deprecated in its description with a pointer to the replacement, and give users a release where both work.
Testing against a real client
Unit tests prove your handlers run, not that a model can use them. Two levels of testing are worth building.
The mechanical level is the MCP Inspector, the reference developer tool: one package with a web UI, a terminal UI and a scriptable CLI. Wire the CLI into CI:
npx @modelcontextprotocol/inspector --cli node path/to/server/index.js --method tools/list
It also calls tools with arguments and emits JSON, which makes schema regressions and broken error paths catchable in CI.
The behavioural level is harder to automate. Connect the server to a real agent β /mcp-servers and pages like /apps/claude-code/mcp-servers show which clients support what β and give it ten realistic tasks. Watch which tool it reaches for first, whether it recovers from a malformed call, and how many turns a task takes. Then delete or merge the tools it never picked. That loop, not the specification, is what makes a server good.
What to do next
- Ask of each tool: would a competent agent reach for this, or is it here because the API has an endpoint for it? Merge or cut.
- Add a read-only mode that removes write tools rather than refusing them, and make it the documented default for production credentials.
- Rewrite every error string to name the field, the rule and the recovery step.
- Decide your result-size policy β default limit, filters, truncation notice β and put it in the tool descriptions.
- Run the Inspector CLI in CI, then do ten real tasks through a real client and prune what the model ignores.
If you package instructions alongside the server as an Agent Skill (/skills), Writing One Agent Skill That Works in Several Apps covers the portability half of the problem. For the review and approval process around a server that touches company data, see /guides/business; for running one as a team of one, /guides/solo-builder. The wider map of what connects to what is at /ecosystem.
Read the official docs forβ¦
- Tool definitions, results and error semantics β the exact fields,
structuredContent,outputSchemaand theisErrorcontract: Tools. - Transports and authorisation β stdio versus Streamable HTTP, OAuth 2.1, protected resource metadata and audience validation: Authorization.
- Threat models β token passthrough, confused deputy, SSRF, state handle hijacking and scope minimisation: Security Best Practices.
- Version compatibility β how versions are declared and how dual-era servers behave: Versioning and Compatibility.
- Tool design research β consolidation, namespacing, token efficiency and how to evaluate a change: Writing effective tools for AI agents, plus the MCP Inspector reference for its flags.
Mentioned in this guide
GitHub MCP Server π MCP serverFree
GitHub's official MCP server: issues, pull requests, code, Actions and security alerts
Supabase MCP π MCP serverFreemium
Manage Supabase projects, databases, migrations and edge functions from your agent
MCP Server for MySQL π MCP serverOpen source
Read-only-by-default MySQL access with opt-in writes, SSH tunnels and PII masking