Back to Blog

Why Your LLM Uses So Many Tokens for Small Tasks in Claude Code

Scott Brooks, Mana FounderScott Brooks, Mana Founder
•
hidden token costsclaude code token usageretry loop overheadtool call overheadverbose response costtoken optimization
Why Your LLM Uses So Many Tokens for Small Tasks in Claude Code

The Hidden Actions Your LLM Takes That Waste Your Token Budget Every Day

You asked for a small change.

Your assistant read half the repo, called several tools, retried twice, and returned a long explanation you did not need.

Now token usage is 10x what you expected.

This is common, and it is fixable.

If you have asked, "why does my LLM use so many tokens for small tasks in Claude Code?", the short answer is hidden overhead from reasoning depth, tool calls, retries, and verbose responses.

We use Claude Code examples because this behavior is easy to observe there, but the same overhead patterns appear across modern AI coding assistants.

Quick answer (30 seconds)

Small prompts can trigger big token bills when assistants over-reason, call too many tools, retry repeatedly, and return verbose outputs. Vague instructions also force broad exploration. The fastest fixes are tighter prompt scope, concise-output instructions, and stopping unproductive retry loops early.


Why does my LLM use so many tokens for small tasks in Claude Code?

Because visible prompt size is only part of total cost. Tool schemas, tool payloads, chain-of-actions, retries, and long responses all add hidden token overhead, especially when repeated across turns.


Hidden cost #1: thinking depth you did not need

Some tasks benefit from deeper reasoning.

Examples:

  • architecture tradeoffs
  • complex debugging
  • multi-file refactors

Some tasks do not.

Examples:

  • simple rename
  • one-line null check
  • formatting cleanup

When high reasoning depth is used on low-complexity work, you overpay.

Match thinking depth to task complexity.


Hidden cost #2: tool call overhead

Each tool call has two token costs:

  1. request framing (tool schema + parameters)
  2. response payload

Single calls may be small. Chains add up quickly:

  • search
  • read
  • search again
  • read again
  • summarize

If responses are verbose, token spend spikes before code changes happen.


Hidden cost #3: retry loops

Retry loops are expensive because each retry carries context history.

Pattern:

  • attempt fails
  • assistant retries with a slightly different approach
  • fails again
  • retries again

You now pay for:

  • original context
  • failed attempts
  • additional outputs

Practical rule: if the same path fails twice, reframe the prompt with tighter constraints instead of letting retries continue.


Hidden cost #4: verbose response syndrome

Most assistants default to being thorough.

That is helpful for teaching, but expensive when you only need execution.

Examples:

  • yes/no answer becomes 300 words
  • code diff is followed by a long narrative summary
  • each step is explained when you only needed the result

Set response style explicitly for routine tasks (for example: concise, no trailing summary).


Hidden cost #5: exploration tokens

When prompts are vague, assistants explore broadly to avoid mistakes.

  • "Fix the bug" -> many files read
  • "Fix null check in src/auth.ts line 42" -> narrow read scope

Specific prompts reduce exploration overhead and usually improve output quality.


Practical optimization checklist

  1. Match model and reasoning depth to task complexity.
  2. Use exact file paths and error locations whenever possible.
  3. Ask for concise output on routine edits.
  4. Stop unproductive retry loops and reframe the prompt.
  5. Keep sessions focused; clear or reset between unrelated tasks.
  6. Reduce tool-response bloat with filtered outputs.

These are small behavior changes with outsized token impact.


Key takeaways

  • Hidden overhead often comes from thinking depth, tool calls, retries, verbosity, and exploration.
  • These hidden actions can multiply true cost far beyond prompt size.
  • Prompt specificity is still the highest-leverage optimization.
  • Better guidance beats more retries.
  • Efficiency is mostly operational, not mystical.

FAQ

Why does my LLM use so many tokens for small tasks in Claude Code?

Small prompts can still trigger expensive workflows behind the scenes, including tool calls, retries, context carryover, and verbose responses. Total token cost is driven by the full execution path, not just prompt length.

What hidden actions increase token usage in AI coding assistants?

The biggest hidden drivers are unnecessary reasoning depth, repeated search/read tool chains, retry loops, verbose output, and broad exploration caused by vague prompts.

Do tool calls and retries significantly raise token costs?

Yes. Each tool call adds request and response tokens, and retries repeat context while adding new outputs. Costs compound quickly in multi-step sessions.

How does verbose output waste tokens?

Long explanations and oversized summaries consume context even when you only need a direct result. That added text increases immediate and downstream token spend.

Why do vague prompts increase token spend?

Vague requests force assistants to search broadly and read more files to reduce ambiguity. Specific paths, errors, and constraints usually reduce exploration overhead.

How can I reduce token waste without lowering answer quality?

Match reasoning depth to task complexity, provide exact scope, request concise outputs for routine edits, and reframe quickly when retries are not making progress.

Do these hidden token costs apply beyond Claude Code?

Yes. Implementations differ, but most agentic coding tools face similar overhead patterns when workflows include tool orchestration, retries, and large response payloads.


Related reading


CTA

If tool-response bloat is the hidden tax in your workflow, Mana filters oversized MCP/tool outputs so your assistant can stay effective without carrying unnecessary context weight.

Join the Waitlist