Back to Blog

Why MCPs Use So Many Tokens in Claude Code (and How to Reduce Waste)

Scott Brooks, Mana FounderScott Brooks, Mana Founder
•
mcp token overheadmodel context protocol costsclaude code mcp tokensdynamic tool loadingresponse filteringtoken optimization
Why MCPs Use So Many Tokens in Claude Code (and How to Reduce Waste)

The Power and Pain of MCPs (Token Waste)

MCP is one of the biggest leaps in AI coding.

It turns your assistant from a text generator into a tool-using agent that can work across files, APIs, and services.

That power is real. So is the token overhead.

This post explains where MCP token cost comes from and how to optimize it without giving up capability.

If you have asked, "why do MCPs use so many tokens in Claude Code?", the short answer is schema loading plus bloated tool responses.

We use Claude Code examples because this overhead is easy to see there, but the same MCP dynamics show up across modern AI coding assistants.

Quick answer (30 seconds)

MCP token usage spikes because models often load large tool schemas and then process oversized tool responses across many turns. Costs compound when too many tools are exposed at once or when responses include low-signal payloads. The biggest wins are task-scoped tool exposure, on-demand schema loading, and response filtering.


Why do MCPs use so many tokens in Claude Code?

Because MCP adds a structured tool layer the model must understand and repeatedly use. That means paying for both schema context and response context. In active coding sessions, response payload bloat often becomes the larger long-run token driver.


Why MCP matters

Without MCP, your model can only read and write text.

With MCP, it can:

  • call external tools
  • inspect systems
  • chain multi-step actions
  • complete workflows that would otherwise require manual copy/paste

This is why MCP feels like a superpower in coding workflows.

The tradeoff is startup and response overhead that many teams underestimate.


Where the token overhead comes from

MCP overhead usually comes from two places.

1) Tool definition schemas

Each MCP tool is described by structured schema text (name, description, parameters, types, validation rules).

When many tools are loaded at once, schema overhead can get large before useful work begins.

Reported real-world examples include:

  • stacks with dozens of tools adding tens of thousands of tokens
  • large API surfaces producing very large schema payloads
  • repeated schema loading across sessions or agents

2) Tool response payloads

Even after startup, response volume can dominate costs.

Examples:

  • browser tools returning full HTML when you only need extracted text
  • file/listing tools returning full metadata when you only need names
  • API discovery responses returning fields you did not ask for

Large responses fill context with low-signal tokens.


The response bloat problem is bigger than most people think

Many teams focus only on startup schema cost.

But response bloat often burns more tokens over a full session.

Why:

  • tool calls happen repeatedly
  • each response can be large
  • outputs stay in context and compound

So even if schema cost is controlled, noisy responses can still crush your budget.


Benchmarks are directionally clear

Across public discussions and tool benchmarks, one pattern keeps showing up:

  • raw MCP workflows can cost much more than direct CLI/API calls for simple operations
  • better tool design (dynamic toolsets, scoped schemas, filtered outputs) can reduce that gap dramatically

Important nuance:

MCP is not bad or too expensive by default.

MCP enables workflows that plain CLI wrappers cannot match.

The real optimization target is how MCP data is loaded and returned.


The optimization path (not removal)

The answer is not "stop using MCP."

The answer is to keep capability while reducing waste:

  1. On-demand schema loading
    Load full tool definitions only when needed for the current task.

  2. Response filtering
    Return the relevant slice of a tool response, not the full dump.

  3. Prompt caching for stable definitions
    Repeated unchanged context can be billed at lower rates on supported providers.

  4. Task-scoped tool selection
    Expose only tool families relevant to the current workflow.

  5. Keep context lean between steps
    Prevent old tool output from dragging through every turn.

This keeps MCP powerful and makes token economics manageable.


Key takeaways

  • MCP is essential for real tool-using AI workflows.
  • MCP token overhead comes from both schema loading and response bloat.
  • Response bloat is often the bigger long-session cost driver.
  • On-demand loading + response filtering preserves capability while cutting waste.
  • Optimize MCP architecture, not MCP away.

FAQ

Why do MCPs use so many tokens in Claude Code?

MCP sessions typically include both tool schema text and tool outputs. When many tools are available and responses are verbose, token usage can rise quickly, especially across multi-step workflows.

What causes MCP token overhead in AI coding tools?

The two biggest drivers are schema overhead (tool definitions, parameters, validation rules) and response overhead (large payloads returned from tools). Repetition across turns compounds both.

Are MCP tool schemas expensive in token usage?

They can be. Large tool catalogs and detailed schemas may consume significant tokens before meaningful task work starts, particularly when loaded too broadly for a narrow task.

Why do MCP responses consume so many tokens?

Tools often return more data than needed, such as full HTML, exhaustive metadata, or oversized logs. That low-signal output still occupies context and increases downstream token spend.

How can I reduce MCP token usage without losing capability?

Use task-scoped tool exposure, load detailed schemas on demand, filter responses to only relevant fields, and avoid carrying stale tool output through every turn.

Is dynamic tool loading better for MCP token efficiency?

In many cases, yes. Loading only the tool families required for the current task reduces startup context and keeps sessions leaner without removing MCP capability.

Do MCP token costs apply beyond Claude Code?

Yes. Implementations differ, but most agentic coding tools that rely on structured tool calls face similar schema and response overhead patterns.


Related reading


CTA

If MCP overhead is eating your budget, Mana helps by loading schemas on demand and filtering bloated tool responses so your model gets what it needs, not everything possible.

Join the Waitlist