TL;DR: MCP tool definitions are recurring context overhead, not free metadata. Count the exact tools/list catalog with your target provider’s token counter, attribute cost by server and tool, then deduplicate schemas, consolidate overlapping tools, defer low-priority catalogs, or route suitable workflows to code execution. Published examples report reductions from 60% to 98.7%; validate selection quality after every change.
Why MCP token waste is a real cost center
MCP token waste is a real cost center because connected tool definitions are serialized into model context alongside task data, and published measurements put their footprint as high as 72% of a 200k-token window, according to the MCP tools specification, Scott Spence’s measurements, and reporting on Perplexity’s decision to step back from MCP.
The cost recurs whenever a request includes the tools catalog: a large catalog can crowd out retrieved documents, conversation history, and task-specific instructions. A nominally large context window does not make that cost irrelevant; it can also increase prefill work and, depending on the serving setup, memory pressure.
The cost is especially easy to miss when teams think in terms of tool calls. A tool that is never selected can still occupy the model’s prompt if its definition is included. Track catalog size separately from invocation count and call-result size. For a related treatment of the caching tradeoffs, see prompt cache hit-rate engineering.
What actually gets sent in a tools/list payload
A tools/list payload contains each tool’s name, description, and inputSchema, with optional outputSchema, annotations, and _meta, according to the MCP tools specification; LeanZero’s 145-tool probe measured the schema at 81.1% of definition tokens, descriptions at 15.7%, and names at 3.2%.
That distribution makes schema structure the first place to inspect, but it does not make descriptions or names free. Long prose, repeated property descriptions, verbose enum values, and overly broad input shapes can all add overhead. Optional fields can matter too if a client forwards them to the model.
Measure the representation your client actually supplies to the model, not just the server’s raw response. A host may transform definitions, add wrappers, or render them through a chat template. LeanZero found its rendered form of 145 tools cost more than the raw serialized definitions, so a raw JSON counter is useful for comparison but not always the final prompt total. See LeanZero’s probe and methodology.
Measure definition tokens before you cut
Measure MCP definitions before changing them: the provider’s count_tokens API is the billing-relevant measure, while a local tokenizer is an offline triage proxy; in MCP issue #2808, an 11-tool catalog ranged from about 100 tokens for a light definition to about 1,000 for a heavy one, or 5–15 times a type-only schema.
The practical sequence is to capture the same catalog and rendered request shape your production client uses, then count with the provider’s API. For fast iteration across many servers, use a local tokenizer to find expensive definitions, but do not treat its output as billed usage. The Anthropic token-counting documentation describes an API-based way to estimate input tokens before sending a request.
Runnable tools/list audit
This script launches one stdio server, performs the MCP initialization handshake, requests tools/list, and reports a per-tool breakdown. Component token counts are diagnostic; tokenizer boundaries mean those component counts need not sum exactly to the whole-definition count.
import argparse, json, shlex, subprocess, sys, tiktoken
# cl100k_base is an offline proxy; billing truth requires the provider's count_tokens API.
enc = tiktoken.get_encoding("cl100k_base")
def count(value):
return len(enc.encode(json.dumps(value, separators=(",", ":"), ensure_ascii=False)))
def send(proc, message):
proc.stdin.write(json.dumps(message) + "\n")
proc.stdin.flush()
def receive(proc, wanted):
while True:
line = proc.stdout.readline()
if not line:
raise RuntimeError("Server exited before responding")
try:
msg = json.loads(line)
except json.JSONDecodeError:
continue
if msg.get("id") == wanted:
if "error" in msg:
raise RuntimeError(msg["error"])
return msg["result"]
ap = argparse.ArgumentParser()
ap.add_argument("--server", required=True)
ap.add_argument("--json", action="store_true")
args = ap.parse_args()
proc = subprocess.Popen(shlex.split(args.server), stdin=subprocess.PIPE,
stdout=subprocess.PIPE, text=True, bufsize=1)
try:
send(proc, {"jsonrpc":"2.0","id":1,"method":"initialize",
"params":{"protocolVersion":"2025-06-18",
"capabilities":{},"clientInfo":{"name":"mcp-audit","version":"1"}}})
receive(proc, 1)
send(proc, {"jsonrpc":"2.0","method":"notifications/initialized","params":{}})
send(proc, {"jsonrpc":"2.0","id":2,"method":"tools/list","params":{}})
tools = receive(proc, 2).get("tools", [])
finally:
proc.terminate()
rows = []
for tool in tools:
parts = {k: count(tool[k]) for k in ("name", "description", "inputSchema") if k in tool}
parts["other"] = count({k:v for k,v in tool.items()
if k not in ("name","description","inputSchema")})
rows.append({"name":tool.get("name",""), "total":count(tool), **parts})
mean_schema = sum(r.get("inputSchema",0) for r in rows) / max(len(rows),1)
for r in rows:
r["flag"] = r.get("inputSchema",0) > 2 * mean_schema
result = {"tool_count":len(rows), "total_tokens":sum(r["total"] for r in rows),
"avg_tokens_per_tool":sum(r["total"] for r in rows)/max(len(rows),1),
"tools":sorted(rows,key=lambda r:r.get("inputSchema",0),reverse=True)}
if args.json:
print(json.dumps(result, indent=2))
else:
print(f"Tools: {result['tool_count']} Definitions: {result['total_tokens']} tokens"
f" Average: {result['avg_tokens_per_tool']:.1f}/tool")
print("Tool | Total | Name | Description | Schema | Other | Flag")
for r in result["tools"]:
print(f"{r['name']} | {r['total']} | {r.get('name',0)} | "
f"{r.get('description',0)} | {r.get('inputSchema',0)} | "
f"{r['other']} | {'>2x mean' if r['flag'] else ''}")
For CI, keep the JSON output as a baseline and alert on material increases per server or tool. It is a measurement aid, not a replacement for checking the provider’s actual request accounting. Pair it with the operational signals in the MCP server observability guide.
Attribute cost per server and per tool
Per-server attribution exposes catalog skew that aggregate counts hide: LeanZero measured GitHub’s 26 tools at 3,562 definition tokens and Notion’s 24 tools at 16,772, a 10.2× spread despite similar tool counts; across its sample, LeanZero reported a correlation of only r = 0.518 between tool count and token cost.
Use the audit output to rank servers by total definitions, then rank tools within each server. Token cost per tool is a useful signal, not a quality score: a large schema might encode a genuinely complex operation, while a short description can still leave ambiguous behavior. Compare definition cost with selection frequency and task coverage before removing anything.
| Server / source | Tool count | Definition tokens | Tokens per tool | Share of 200k window |
|---|---|---|---|---|
| mcp-omnisearch before (Scott Spence) | 20 | 14,214 | 710.7 | 7.1% |
| mcp-omnisearch after consolidation (Scott Spence) | 8 | 5,663 | 707.9 | 2.8% |
| GitHub (LeanZero) | 26 | 3,562 | 137.0 | 1.8% |
| Notion (LeanZero) | 24 | 16,772 | 698.8 | 8.4% |
| 10 real servers, all 145 tools (LeanZero) | 145 | 40,784 | 281.3 | 20.4% |
LeanZero reports that rendering those 145 tools through its chat template raises the total to 51,530 prompt tokens (25.8% of 200k), reserving 3.16 GiB of KV cache. For alternative local inspection workflows, compare mcp-token-analyzer with the script above; their counting assumptions should still be checked against the target provider.
Cut 1 — deduplicate repeated JSON Schema $defs
Deduplicating repeated JSON Schema $defs can remove large amounts of repeated text: LeanZero found nine identical definitions repeated across Notion’s 24 tools, accounting for 30.1% of the payload, and reports schema tokens falling 77.7%, from 15,780 to 3,516, after stripping them.
Audit identical $defs across tools and servers before rewriting schemas. Deduplication is most effective when the client or provider can preserve shared references; if a downstream layer expands references before prompt construction, measure that rendered result too. Avoid changing validation semantics merely to reduce bytes.
The SEP-1576 proposal identifies repeated fields as a potential standardization target: its review found owner in 36 of 60 schemas, repo in 39 of 60, and required in nine of 60. Those counts support investigating common structures, not blindly removing them. Keep schemas explicit enough for reliable selection and valid arguments.
Cut 2 — consolidate tools and tighten schemas
Consolidating overlapping tools can reduce definition cost without shrinking useful capability: Scott Spence reports that mcp-omnisearch fell from 20 tools and 14,214 tokens to eight tools and 5,663 tokens, a 60% reduction; Pydantic’s engineering guidance recommends snapshot-testing responses and returning useful Markdown rather than raw JSON.
Look for tools that differ only by a fixed endpoint, mode, or resource type. A single well-designed operation with a constrained enum may replace several near-duplicates, provided the model can still express the needed action clearly. Then shorten schema descriptions to the information needed for correct selection and argument construction.
Schema efficiency is not just input size. If a tool returns a large, repetitive JSON object when the agent needs a concise summary or a few fields, output handling can dominate the full interaction. Snapshot-test representative responses when changing output shape, and preserve the data needed for downstream reasoning. Use the MCP servers repository as a reference for real server catalogs, not as a guarantee that any particular schema is token-efficient.
Cut 3 — defer or load tools on demand
Deferring tools can keep rarely used definitions out of the planning context: ProMCP, in its ACL 2026 study, reports custom configurations consuming 52–63% of planning tokens with 169 resident tools, versus 2.1% with deferred configuration; SEP-1576 also discusses embedding-based tool retrieval as a way to select a smaller relevant catalog.
Treat deferred loading as a routing problem. A lightweight first-stage selector can identify likely server or tool families, then expose only their definitions for the current task. Measure both the resident catalog and the definitions introduced after selection, and test retrieval misses: an omitted tool cannot be chosen later unless your system has a reliable way to fetch it.
The 2026-07-28 MCP release and its specification changelog add cache hints and deterministic list ordering, which can improve stability for caching. They do not reduce catalog size: every definition still costs context when included. Read the 2026-07-28 stateless MCP spec with that distinction in mind.
Cut 4 — route chained work to code execution
Code execution can sharply reduce repeated tool and intermediate-data context for suitable chained workflows: Anthropic reports a reduction from 150,000 to 2,000 tokens, or 98.7%, while AIMultiple’s test reports input tokens falling 78.5% and total tokens 77.4%, with output tokens up 120%, 100% success in both conditions, and latency changing from 9.66s to 10.37s.
These results are workload-specific, not a blanket argument to replace tools. The pattern is most promising when a task needs many sequential operations, filtering, or aggregation and only a compact result must reach the model. Sarah Deaton reports 50,000–73,000 tokens reduced to 9,500–10,000, or 80–87%, in her MCP code-execution examples.
Compare total tokens, latency, success, and the quality of the final result on your own representative tasks. Code execution also changes the operational and security boundary, so it needs appropriate controls. This is not a build tutorial; see our code-execution MCP server build for that separate implementation topic.
Placement policy — decide what is always resident
A placement policy decides which tool definitions are always resident, task-scoped, manually enabled, split into smaller catalogs, shrunk, or rejected; Vorp Labs’ overhead worksheet separates catalog overhead from response overhead, while the mcp-token-audit manifest demonstrates deferred loading.
Set a budget per agent role, not just a global maximum. Keep frequently selected, broadly useful definitions resident; load specialized tools only when the task warrants them; and reject catalogs whose value does not justify their cost. A split catalog can also keep unrelated domains from competing during tool selection.
The mcp-token-audit project gives a worked example in which a 3,096-token manifest is 1.5% of a 200k window, and a five-server setup falls from 77K to 8.7K tokens with deferred loading. Treat these as that project’s example, not a guaranteed outcome. Track catalog overhead separately from response overhead, because optimizing definitions will not solve oversized tool results.
When to skip MCP entirely
Skipping MCP can be the right choice when a catalog’s context cost outweighs the value of its integration: reporting on Perplexity attributes 143,000 of a 200,000-token window, or 72%, to tool definitions, while an independent fact-check cautions that the broader claim of an organization-wide abandonment is not independently established.
Do not generalize one reported decision into a universal protocol verdict. Verify the attribution and deployment context, then compare MCP with direct APIs, narrower adapters, or a task-specific integration. If a server contributes little unique capability and imposes persistent catalog overhead, removing it may be better than optimizing its schemas.
The decision should include maintenance, permissions, reliability, and the value of a consistent tool interface—not just context tokens. Keep the integration only when its operational benefits justify its cost for the agent’s real workload. For a broader view of the tradeoffs, use the agent cost-optimization playbook.
The Bottom Line
The operational verdict is to audit tools/list with a real provider counter, attribute cost per server and per tool, then apply deduplication, consolidation, deferred loading, or code execution; published examples from Scott Spence, LeanZero, Anthropic, and AIMultiple show 60–98.7% reductions, but production teams must validate selection quality after shrinking catalogs.
Start with the highest-cost server, not the easiest one to edit. Capture a baseline, make one change at a time, and compare definition tokens, task success, tool-selection accuracy, latency, and output tokens on a representative evaluation set. A smaller catalog that causes missed tools or invalid calls is not a successful optimization.
Keep the resulting audit in CI or a periodic review so server upgrades do not silently restore expensive definitions. Use MCP testing and CI/CD to make schema changes and behavior checks part of the same release process.
How This Guide Was Built
This guide separates published measurements from operational recommendations: its figures come from linked studies, project measurements, issue discussions, and official documentation, while the suggested audit workflow is a client-side procedure rather than a claim of hands-on benchmarking.
This is desk research over published measurements and official docs — no MCP servers were run hands-on by the author.
Primary sources consulted include the MCP tools specification, MCP issue #2808, SEP-1576, ProMCP’s ACL 2026 paper, LeanZero’s measurements, Scott Spence’s report, Anthropic Engineering, AIMultiple, Sarah Deaton, Vorp Labs, and Pydantic. The Perplexity discussion also draws on AgentMarketCap and ByteIota, alongside the independent fact-check linked above.
FAQ
This FAQ answers five operational questions about MCP definition overhead using provider counting guidance, published token measurements, and the cited code-execution examples; the key distinction is between the cost of definitions included in context and the separate costs of tool outputs, orchestration, and execution.
Are MCP tool definitions really the biggest context cost?
Sometimes, but not in every workload. Perplexity-related reporting attributes 143,000 of a 200,000-token window to tool definitions, while LeanZero measured 40,784 raw definition tokens for 145 tools before chat-template rendering. Compare the catalog with conversation history, retrieved content, and tool results in your own requests.
Does prompt caching fix the token-waste problem?
No. Caching may reduce repeated processing or cost under a provider’s pricing rules, but it does not make an oversized catalog smaller or create room for task context. The 2026-07-28 MCP changelog improves ordering and cache stability; it does not reduce the number of definitions sent.
Is code execution always better than MCP tools?
No. Code execution is most compelling for workflows with chained operations or large intermediate results, not every discrete action. Anthropic’s published example reports 150,000 tokens reduced to 2,000, but that is a workload-specific result. Compare reliability, latency, security controls, total tokens, and task quality before routing production work.
How accurate are tiktoken-based token counts for MCP auditing?
They are useful for ranking tools and comparing changes under a consistent local setup, but they are not billing truth. Tokenization and request rendering can vary by model and provider. The Anthropic token-counting documentation describes provider-side counting; use that for billing-relevant estimates.
What should I cut first in a production MCP setup?
Start with the server or tools contributing the most definition tokens, then inspect repeated schemas, overlapping operations, and rarely needed catalogs. LeanZero’s Notion probe found repeated $defs that were candidates for deduplication. Preserve a baseline and validate tool selection and task success after every change.
📖 Related Reads
- ToolBrain — tool reviews, LLM comparisons, and AI workflow guides
- Hermes Tutorials — Hermes Agent setup, configuration, and advanced workflows
- CodeIntel Log — code quality, debugging, and software engineering benchmarks
Cross-links automatically generated from NiteAgent.
← Back to all posts


