Back to Blog

Why Claude Code Uses So Many Tokens Before Your First Prompt

Scott Brooks, Mana FounderScott Brooks, Mana Founder
•
claude code startup tokensllm startup token costsystem prompt overheadmcp schema loadingskills token overheadon-demand schema loadingtoken optimization
Why Claude Code Uses So Many Tokens Before Your First Prompt

Saving Tokens at Conversation Start with LLMs

Why usage can spike before your first real prompt in AI coding workflows

You're in Claude Code, you type "fix the bug on line 42," and usage jumps faster than expected.

That is usually real, not a dashboard glitch. A meaningful chunk of token spend can happen before your first serious prompt lands.

This post breaks down what loads at startup, why that overhead exists, and how to keep capability while cutting waste.

If you have asked, "why does Claude Code use so many tokens before first prompt?", startup context loading is usually the reason.

We'll use Claude Code examples because this pattern is easy to spot there, but the same startup token mechanics show up across modern AI coding assistants.

Quick answer (30 seconds)

Startup token cost comes from everything loaded before your task gets moving: system instructions, project rules, MCP schemas, skills, and carried context. In capable setups, this baseline can be large before your first meaningful prompt. The goal is not to remove capability, but to load heavy context only when it is needed.


Why does Claude Code use so many tokens before first prompt?

Because each session may preload multiple layers of context and tooling before real task execution begins. Those layers are useful, but they still consume tokens. In advanced setups, startup overhead can rival or exceed the cost of several normal turns.


What loads before your first message

Most modern coding setups preload several layers so your assistant can act like an agent, not just a chatbot.

Typical startup layers:

  1. System prompt

    • Base behavior and safety instructions
    • Often around 1,000 to 3,000 tokens
  2. Project rules files (for example, CLAUDE.md and related docs)

    • Coding standards, workflows, repo conventions
    • Commonly 500 to 10,000+ tokens depending on project maturity
  3. MCP tool definitions

    • Schemas describing each connected tool, parameters, and constraints
    • Often 5,000 to 55,000+ tokens depending on number and complexity of servers/tools
  4. Skills/persona instruction sets

    • Domain-specific instructions for specialized tasks
    • Can be small, but can also become large when multiple skills are always active

Add those layers up and startup overhead in the 20K to 70K range can be normal for a capable setup.

That number sounds high. It is also why your assistant can do useful work across files, tools, and workflows.


The cost of capability (and why this is not just waste)

It is easy to look at startup tokens and think, "this is all overhead."

Some of it is overhead. Some of it is capability you are intentionally buying.

  • Without strong project instructions, output quality gets inconsistent.
  • Without tool schemas, the model cannot reliably call external tools.
  • Without specialized skill instructions, it can miss domain-specific nuance.

So the better question is not, "should I remove all this?"

It is, "how do I load this efficiently so I pay for what I use?"


What MCP tool loading actually costs

MCP is one of the most useful upgrades in modern AI coding workflows. It is also one of the largest startup token contributors.

Why: each tool comes with a schema, and schema text adds up quickly.

Reported examples from community and vendor discussions include:

  • A stack with GitHub + Slack + Sentry (~40 tools) around 55,000 tokens of schema overhead
  • Some individual MCP servers adding roughly 18,000 tokens per message context in certain setups
  • Large surface-area APIs (like full platform APIs via MCP) creating very large schema payloads in worst cases

Prompt caching can reduce repeated schema cost when content stays stable.

But caching is not your only lever.

A stronger pattern is on-demand loading:

  • discover what servers/tools are available with low-cost metadata first
  • load full schema only when the model actually needs that specific tool

Same capability. Less upfront bloat.


Skills loading impact

Skills are powerful. They also have startup weight.

When many skills are always active, full instruction sets can inflate baseline token usage before real work starts.

Reported issue threads in the ecosystem have shown:

  • 50K+ tokens of startup contribution from skills in heavy configurations
  • meaningful baseline reductions when lazy-loading patterns are used instead of always-on loading

The same pattern as MCP optimization applies:

  • keep high-value foundational guidance always available
  • load specialized depth only when task intent requires it

That gives better budget efficiency without dumbing down the assistant.


Practical ways to reduce startup token cost

You do not need to neuter your setup to improve efficiency.

Use these five moves first:

  1. Keep root instruction files concise Put universal rules at root, move deep details into scoped files.

  2. Split project guidance by feature area Load heavy instructions only when working in that area.

  3. Audit always-on skills quarterly If a skill is rarely used, do not keep it permanently active.

  4. Stabilize repeated instruction blocks Stable text increases cache hit rates and lowers recurring cost.

  5. Prefer on-demand schema loading for MCP Pull full definitions only for tools needed in the current task.

If you do only these five, you can usually reduce startup tax while keeping output quality steady.


Key takeaways

  • Startup overhead is real and often large in capable LLM coding environments.
  • A 20K to 70K startup range can be normal once rules, tools, and skills are loaded.
  • This overhead buys real capability, so elimination is usually the wrong strategy.
  • The best optimization is efficient loading: scoped rules, selective skills, on-demand schemas, and caching.
  • You can keep power and still cut token waste.

FAQ

Why does Claude Code use tokens before I type anything?

Because the environment can preload system instructions, project rules, tool schemas, active skills, and prior context. Those tokens are part of startup capability and are counted before your first meaningful prompt.

What contributes most to LLM startup token cost?

Usually a mix of long instruction files, MCP tool schemas, and always-on skills. The relative share depends on your setup, but MCP-heavy environments often see schema overhead as a major contributor.

Is 20K-70K startup token usage normal?

In capable coding-agent setups, yes, it can be normal. Systems with rich tooling and guidance can consume significant baseline tokens before task-specific prompts begin.

Are MCP schemas a major source of startup overhead?

They often are. Schema definitions for many tools can add substantial token volume, especially when full definitions are loaded up front rather than on demand.

Can prompt caching reduce startup token cost?

Yes, for stable repeated content. Caching can reduce recurring costs on eligible repeated prompt segments, but it does not fully solve overhead from constantly changing context.

How do I reduce startup token usage without losing capability?

Keep root instructions concise, scope deep guidance by area, audit always-on skills, stabilize reusable text for better cache hit rates, and load full MCP schemas only when needed.

Do startup token costs apply to Cursor, Cline, and other AI coding tools?

Yes. While implementations differ, the same pattern appears across tools that preload instructions, context, and tool metadata.


Related reading


CTA

If your usage spikes before real work starts, you do not need to remove tools. You need smarter loading.

Mana reduces startup bloat by loading MCP schemas on demand and filtering tool-output noise, so you keep capability while extending useful coding capacity.

Join the Waitlist