Engineering

Why a long Claude Code session costs more on every message

Editorial · Reveneau · October 5, 2026

Why a long Claude Code session costs more on every message

Claude Code is Anthropic's coding tool. You type a request in a terminal, the text window where you type commands, and an AI model reads files, runs commands and edits code for you. All of that work is counted in tokens. A token is a piece of text that the model processes, and Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters, or 0.75 of an English word.

The model keeps nothing between two requests. Every message you send becomes a request that contains the whole conversation again. This post explains what that does to a long session, and then gives the two habits Anthropic's own documentation names first. Both are free.

Every request carries everything that came before it

Anthropic's documentation on prompt caching, read on 4 October 2026, says Claude Code sends the full context with every message: "the system prompt, your project context, every prior message and tool result, and your new message". The place that holds all of this is the context window, which is all the text the model reads in one request. Your new line goes at the end.

One message from you is often several requests to the model. Each time Claude reads a file or runs a command, Claude Code sends another request that contains the result together with the whole conversation. The rule is Anthropic's: "Token costs scale with context size: the more context Claude processes, the more tokens you use."

Anthropic's own simulation of a session shows the proportions. Seven items load before the user types anything, and by our arithmetic on Anthropic's figures they add up to 7,850 tokens. The user's prompt is 45 tokens. The four files Claude reads to answer it add 6,900 tokens, and one run of the test suite adds 1,200 tokens while the screen shows a short status line. Anthropic calls these representative token counts, so they are an illustration. You typed 45 tokens, and the request grew by thousands.

What this does to a session that stays open all day

A session that stays open across unrelated tasks carries the text of all of them in every request. Anthropic's documentation warns that such a session "can use far more of your plan limits than your activity suggests". Its example is a one-line question in a session that has been open all day: it is still counted as usage for the whole conversation.

Two mechanisms lower the price of this repeated reading, and Claude Code runs both without being asked. The first is the prompt cache, a store of request text that the service has already processed. Text read again from the cache is billed at a lower price: on Claude Sonnet 5.5, Anthropic's pricing page lists a cache read at $0.20 per million tokens against $2 for new input. The second is compaction, which replaces a long history with a summary when the context window is close to full.

The cache lowers the price of each token and leaves the count unchanged. An invented example, with arithmetic on Anthropic's list prices read on 4 October 2026, shows the size of the effect. Suppose a conversation has grown to 400,000 tokens on Claude Sonnet 4.6, which Anthropic lists at $0.30 per million tokens for a cache read. One request then costs $0.12 for the repeated reading alone, and a task in which Claude uses tools 20 times sends 20 requests, which is $2.40 before any new work is counted. The cache also expires when it is unused for one hour on a subscription, or five minutes on usage credits, a pay-per-token key or a cloud provider, by Anthropic's documentation. The first message after a longer break is processed in full: the same 400,000 tokens written to the five-minute cache at the listed $3.75 per million tokens cost $1.50, which is 12.5 times the cost of the request that read from the cache.

The page on why Claude Code usage keeps rising in a long session lists eight causes from Anthropic's documentation. Five of them are requests that the session sends while you are away.

Habit one: clear the conversation between unrelated tasks

Anthropic's documentation on managing costs names clearing between unrelated tasks as one of the two habits with the highest impact, and says the /clear command itself costs nothing. The same page names "long sessions that were never cleared" as one of the two usual causes of unexpectedly high spending on an API or cloud provider plan.

/clear starts a new conversation with an empty context. If you want to find the old conversation later, run /rename first, and /resume brings it back.

Two other commands shorten a conversation. /compact replaces the history with a summary, so that the same task can continue with a smaller request. Anthropic's documentation adds a cost: "/compact reads the conversation it summarizes, so compacting a large context is itself a large request." /rewind returns to an earlier prompt, for when the last few steps were a mistake. The page on /clear, /compact or /rewind compares what each one costs and keeps. The rule we recommend: /compact inside one long task, /clear between tasks, /rewind when the work since an earlier prompt should be thrown away.

Habit two: match the model to the job, before the first message

The model sets the price of every token. Anthropic's documentation says Sonnet "handles most coding tasks well and costs less than Opus", and advises keeping Opus for "complex architectural decisions or multi-step reasoning". The model you get when you choose nothing is Opus 5.5, on Pro, Max, Team and Enterprise plans and on the Anthropic API, by Anthropic's model configuration page. "Opus left as the default model" is the other usual cause of unexpectedly high spending that Anthropic names.

The price difference is in Anthropic's list. Claude Opus 5.5 is listed at $4 per million input tokens and $20 per million output tokens. Claude Sonnet 5.5 is listed at $2 and $10. A cache read is $0.20 on both. The saving starts with one command, /model sonnet.

The timing matters as much as the choice. Anthropic's caching documentation says each model has its own cache, so after a switch with /model "the next request reads the entire conversation history with no cache hits, even though the content is identical". A switch in the first minute of a session costs almost nothing. The same switch after a long day processes the whole day again at the new model's price. If a task turns out to need Opus, the cheapest moment to switch is after /clear. The page on which model and effort level to use gives the list prices for every current model and a table of tasks.

What to do first

Run /usage in the session you have open and write down the token counts. On a Pro, Max, Team or Enterprise plan, Anthropic's documentation says the screen flags a behaviour such as long context when it accounts for 10 percent or more of your recent usage. Adopt the two habits for a week, run /usage again, and compare.

Reveneau is an AI software development consultancy. All of its code is written by AI, so token use is a running cost of every Reveneau build, and the guide on how to reduce Claude Code token usage collects every change we recommend. Reveneau is independent of Anthropic, and every figure in this post is Anthropic's own statement about its own product, read on 4 October 2026. Anthropic gives no figure for how much either habit saves, and we give none either. Measure your own session, change one thing, and measure again.

The message you type is the smallest part of the request. Everything before it is the cost.

Sources

Common questions

Why does a long Claude Code session cost more per message?

A long Claude Code session costs more per message because every request contains the whole conversation. Anthropic's documentation, read on 4 October 2026, says Claude Code sends the system prompt, your project context, every earlier message and tool result, and your new message each time. A one-line question in a session open all day is therefore counted as usage for the whole conversation, and each file Claude reads makes the next request larger again.

What does Anthropic name as the two habits with the highest impact on Claude Code usage?

Anthropic's documentation on managing costs, read on 4 October 2026, names clearing between unrelated tasks and matching the model to the task. The same page names long sessions that were never cleared, and Opus left as the default model, as the usual causes of unexpectedly high spending on an API or cloud provider plan. Both habits are free, and /clear itself costs nothing, which is why we put them before any setting.

How often should I run /clear in Claude Code?

Run /clear each time you start work that is unrelated to the current conversation. Anthropic's documentation, read on 4 October 2026, says /clear costs nothing and that old context is sent again with every later message, so the right moment is the change of task. Run /rename first if you want to find the old conversation later, and /resume brings it back.

What is the difference between /clear and /compact in Claude Code?

/clear starts a new conversation with an empty context, and /compact replaces the history with a summary so that the same task can continue. Anthropic's documentation, read on 4 October 2026, says /compact reads the conversation it summarises, so compacting a large context is itself a large request, and that when you want a fresh start /clear costs nothing. Use /compact inside one long task and /clear between tasks.

Which model should I pick before starting a Claude Code task?

Pick Sonnet for most coding tasks and keep Opus for architecture decisions and reasoning across many steps. That is Anthropic's own advice in documentation read on 4 October 2026. The default model on Pro, Max, Team and Enterprise plans and on the Anthropic API is Opus 5.5, which Anthropic's pricing page lists at $4 per million input tokens and $20 per million output tokens, twice the Sonnet 5.5 list price.

Why should I choose the model at the start of a session instead of switching later?

Choose the model at the start because each model has its own prompt cache. Anthropic's documentation, read on 4 October 2026, says that after a switch with /model the next request reads the entire conversation history with no cache hits, even though the content is identical. The cache is the store of request text already processed, and text read from it is billed at a lower price, so a switch late in a long session costs the most.

Does a long session use a Pro or Max plan limit faster?

Yes. Anthropic's documentation, read on 4 October 2026, says a session that has been open for hours can use far more of your plan limits than your activity suggests, because usage counts the whole conversation on every request. On a paid plan the /usage command flags a behaviour such as long context when it accounts for 10 percent or more of your recent usage, and shows a tip beside the flag.

How much does a long session cost in dollars?

Anthropic gives no figure for a long session, so any dollar figure is an invented example on its list prices. In one such example, a conversation of 400,000 tokens on Claude Sonnet 4.6 costs $0.12 per request in cache reads at Anthropic's listed $0.30 per million tokens, read on 4 October 2026, and a task with 20 tool uses sends 20 requests, which is $2.40 before any new work. The example shows the shape, and your own session will differ.

What does the first message after a lunch break cost in Claude Code?

The first message after a break longer than the cache lifetime is processed in full, because the stored text has expired. Anthropic's documentation, read on 4 October 2026, gives the lifetime as one hour on a subscription and five minutes on usage credits, a pay-per-token key or a cloud provider. If the old conversation is no longer needed, run /clear after the break so that the next request has little to process.