Back to Blog

Why Long Claude Code Sessions Get Expensive: Context Window Cost Explained

Scott Brooks, Mana FounderScott Brooks, Mana Founder
•
context windowtoken costclaude code contextclaude code token usageprompt cachinglost in the middlellm token optimizationai coding assistant cost
Why Long Claude Code Sessions Get Expensive: Context Window Cost Explained

Context Window and Its Effect on Token Usage

How context size drives cost, speed, and quality in AI coding sessions

Think of your context window as the model's working memory for this turn.

Everything the model needs right now has to fit in that window.

Every token in that window costs money (or burns usage capacity), whether it helps your task or not.

Our intro to tokens explained what tokens are. This post explains why costs can snowball even when your prompts stay short.

If you have asked, "why do long Claude Code sessions get expensive so fast?", context growth is usually the core reason.

We'll use Claude Code examples because the pattern is easiest to spot there. The same mechanics apply across modern AI coding assistants.

Quick answer (30 seconds)

Context window cost grows because each turn can include system instructions, rules files, tool schemas, prior conversation, file snippets, and the new prompt. In long coding sessions, compounded history can dominate total spend. Keeping context lean usually improves both cost and model quality.


Why do long Claude Code sessions get expensive so fast?

Because each turn can carry much more than your latest prompt: prior conversation, system instructions, rules files, tool schemas, attached files, and generated output. As this stack compounds, token cost rises even when your newest message is short.

What is a context window?

A context window is the maximum number of tokens a model can process in one request.

If a request goes over that limit, the system has to truncate, summarize, compact, or fail.

A larger window gives you more room, but it is not free:

  • more room can mean more irrelevant context
  • more irrelevant context can lower answer quality
  • larger requests cost more tokens each turn

Approximate maximum windows by model family (these change over time):

  • Claude Sonnet: up to 200K tokens
  • GPT-4o: up to 128K tokens
  • Gemini 2.5: up to 1M tokens

The key point: max context is a ceiling, not a target.


How your context window fills up

Most developers focus on the prompt they typed.

The model sees a lot more than that.

Typical layers in a coding workflow:

  1. System prompt (base behavior and policies)
  2. Project rules/instructions (global and folder-level)
  3. Tool definitions/schemas (especially in MCP-heavy setups)
  4. Conversation history (all prior turns still in scope)
  5. Working context (files, code, logs, diffs, error output)
  6. Your latest message

That stack is what makes modern coding agents powerful.

The goal is not to remove capability. The goal is to make each layer intentional.


The compounding cost of conversation

The expensive part is not one prompt. It is repetition.

Many systems resend most or all in-scope history every turn.

A simple pattern:

  • Turn 1: small context, low cost
  • Turn 10: history plus tools plus files, moderate cost
  • Turn 30: large history and repeated scaffolding, high cost

So even if your latest prompt is 20 tokens, the request can still be tens of thousands of tokens.

Practical session habits that help:

  • compact or summarize when a thread gets long
  • start a fresh thread after major milestones (for example, after a merged PR)
  • keep sessions task-focused instead of mixing unrelated work
  • avoid attaching broad file sets when one file path would do

Lean context means better quality, not just lower cost

There is a known issue often called "lost in the middle": models can miss important details buried inside long contexts.

When context gets noisy, quality can drop even while cost rises.

Common signs your context is bloated:

  • the model ignores constraints you already gave
  • it repeats earlier mistakes
  • it asks for info you already shared
  • answers get generic even when your task is specific

How to improve quality fast:

  • put critical constraints near your latest prompt
  • include only files/log lines needed for this step
  • prune stale context from old branches or old decisions
  • restate exact success criteria before execution

A lean window is one of the few real win-wins in LLM usage: better output and lower spend.


Cached vs uncached tokens

Not all input tokens are billed equally.

Many providers discount repeated prompt segments through caching.

A common pattern is heavy discounting on cached tokens (for example, around 90% on eligible repeated input for some model tiers).

What usually benefits from caching:

  • stable system prompts
  • stable tool definitions
  • repeated instruction blocks

What often breaks cache efficiency:

  • frequent edits to long instruction text
  • constantly changing large context blocks
  • cache expiry between sessions

Caching does not solve everything, but it can materially reduce repeated overhead when your base context stays stable.


5 ways to keep your context window efficient

  1. Scope each request tightly
    Ask for one concrete outcome per turn.

  2. Front-load only critical context
    Include the exact files and constraints needed for this step.

  3. Trim stale history aggressively
    If context no longer affects the current task, remove or reset it.

  4. Stabilize reusable instructions
    Keep core rules consistent to improve cache hits.

  5. Control tool output volume
    Avoid dumping full logs or huge directory trees when summaries are enough.


Key takeaways

  • Context window cost includes far more than your typed prompt.
  • Conversation history compounds across turns and can become the dominant cost.
  • Bigger context does not automatically mean better results.
  • Lean context improves quality and cost at the same time.
  • Prompt caching can cut repeated overhead when your base context stays stable.

FAQ

What is a context window in an LLM?

A context window is the total token budget the model can process in one request. It includes system instructions, prior messages, tool definitions, referenced files, and your latest prompt.

How does context window size affect token cost?

Larger in-scope context means more tokens processed per turn. If history and tool overhead keep growing, each new request costs more even when your latest prompt is short.

Why do long Claude Code sessions get expensive so fast?

Because history compounds. Each turn can include earlier conversation plus instructions, schemas, and fresh outputs. Over time, the payload grows and token usage accelerates.

Does a larger context window improve answer quality?

Not always. A larger window gives room, but quality can drop when relevant details are buried in noise. Focused context usually produces better answers.

What is the "lost in the middle" problem?

It is the tendency for models to miss or underweight important details buried in the middle of long contexts. This is one reason long, noisy prompts can perform worse.

What are cached vs uncached tokens?

Cached tokens are repeated prompt segments eligible for discounted billing on supported providers. Uncached tokens are newly processed content billed at the standard rate.

How can I reduce context window token waste without hurting coding quality?

Use narrow task prompts, include only relevant files/log lines, reset stale threads, stabilize reusable instructions, and keep tool responses focused. The goal is high-signal context, not minimal context.


Related reading


CTA

If this post helped you understand context window cost, Mana can help you reduce it without giving up performance.

Mana cuts token waste from tool-output bloat and heavy MCP overhead so you get more useful coding sessions from the same budget.

Join the Waitlist