Why a long Claude Code session costs more on every message

Claude Code is Anthropic's coding tool. You type a request in a terminal, the text window where you type commands, and an AI model reads files, runs commands and edits code for you. All of that work is counted in tokens. A token is a piece of text that the model processes, and Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters, or 0.75 of an English word.
The model keeps nothing between two requests. Every message you send becomes a request that contains the whole conversation again. This post explains what that does to a long session, and then gives the two habits Anthropic's own documentation names first. Both are free.
Every request carries everything that came before it
Anthropic's documentation on prompt caching, read on 4 October 2026, says Claude Code sends the full context with every message: "the system prompt, your project context, every prior message and tool result, and your new message". The place that holds all of this is the context window, which is all the text the model reads in one request. Your new line goes at the end.
One message from you is often several requests to the model. Each time Claude reads a file or runs a command, Claude Code sends another request that contains the result together with the whole conversation. The rule is Anthropic's: "Token costs scale with context size: the more context Claude processes, the more tokens you use."
Anthropic's own simulation of a session shows the proportions. Seven items load before the user types anything, and by our arithmetic on Anthropic's figures they add up to 7,850 tokens. The user's prompt is 45 tokens. The four files Claude reads to answer it add 6,900 tokens, and one run of the test suite adds 1,200 tokens while the screen shows a short status line. Anthropic calls these representative token counts, so they are an illustration. You typed 45 tokens, and the request grew by thousands.
What this does to a session that stays open all day
A session that stays open across unrelated tasks carries the text of all of them in every request. Anthropic's documentation warns that such a session "can use far more of your plan limits than your activity suggests". Its example is a one-line question in a session that has been open all day: it is still counted as usage for the whole conversation.
Two mechanisms lower the price of this repeated reading, and Claude Code runs both without being asked. The first is the prompt cache, a store of request text that the service has already processed. Text read again from the cache is billed at a lower price: on Claude Sonnet 5.5, Anthropic's pricing page lists a cache read at $0.20 per million tokens against $2 for new input. The second is compaction, which replaces a long history with a summary when the context window is close to full.
The cache lowers the price of each token and leaves the count unchanged. An invented example, with arithmetic on Anthropic's list prices read on 4 October 2026, shows the size of the effect. Suppose a conversation has grown to 400,000 tokens on Claude Sonnet 4.6, which Anthropic lists at $0.30 per million tokens for a cache read. One request then costs $0.12 for the repeated reading alone, and a task in which Claude uses tools 20 times sends 20 requests, which is $2.40 before any new work is counted. The cache also expires when it is unused for one hour on a subscription, or five minutes on usage credits, a pay-per-token key or a cloud provider, by Anthropic's documentation. The first message after a longer break is processed in full: the same 400,000 tokens written to the five-minute cache at the listed $3.75 per million tokens cost $1.50, which is 12.5 times the cost of the request that read from the cache.
The page on why Claude Code usage keeps rising in a long session lists eight causes from Anthropic's documentation. Five of them are requests that the session sends while you are away.
Habit one: clear the conversation between unrelated tasks
Anthropic's documentation on managing costs names clearing between unrelated tasks as one of the two habits with the highest impact, and says the /clear command itself costs nothing. The same page names "long sessions that were never cleared" as one of the two usual causes of unexpectedly high spending on an API or cloud provider plan.
/clear starts a new conversation with an empty context. If you want to find the old conversation later, run /rename first, and /resume brings it back.
Two other commands shorten a conversation. /compact replaces the history with a summary, so that the same task can continue with a smaller request. Anthropic's documentation adds a cost: "/compact reads the conversation it summarizes, so compacting a large context is itself a large request." /rewind returns to an earlier prompt, for when the last few steps were a mistake. The page on /clear, /compact or /rewind compares what each one costs and keeps. The rule we recommend: /compact inside one long task, /clear between tasks, /rewind when the work since an earlier prompt should be thrown away.
Habit two: match the model to the job, before the first message
The model sets the price of every token. Anthropic's documentation says Sonnet "handles most coding tasks well and costs less than Opus", and advises keeping Opus for "complex architectural decisions or multi-step reasoning". The model you get when you choose nothing is Opus 5.5, on Pro, Max, Team and Enterprise plans and on the Anthropic API, by Anthropic's model configuration page. "Opus left as the default model" is the other usual cause of unexpectedly high spending that Anthropic names.
The price difference is in Anthropic's list. Claude Opus 5.5 is listed at $4 per million input tokens and $20 per million output tokens. Claude Sonnet 5.5 is listed at $2 and $10. A cache read is $0.20 on both. The saving starts with one command, /model sonnet.
The timing matters as much as the choice. Anthropic's caching documentation says each model has its own cache, so after a switch with /model "the next request reads the entire conversation history with no cache hits, even though the content is identical". A switch in the first minute of a session costs almost nothing. The same switch after a long day processes the whole day again at the new model's price. If a task turns out to need Opus, the cheapest moment to switch is after /clear. The page on which model and effort level to use gives the list prices for every current model and a table of tasks.
What to do first
Run /usage in the session you have open and write down the token counts. On a Pro, Max, Team or Enterprise plan, Anthropic's documentation says the screen flags a behaviour such as long context when it accounts for 10 percent or more of your recent usage. Adopt the two habits for a week, run /usage again, and compare.
Reveneau is an AI software development consultancy. All of its code is written by AI, so token use is a running cost of every Reveneau build, and the guide on how to reduce Claude Code token usage collects every change we recommend. Reveneau is independent of Anthropic, and every figure in this post is Anthropic's own statement about its own product, read on 4 October 2026. Anthropic gives no figure for how much either habit saves, and we give none either. Measure your own session, change one thing, and measure again.
The message you type is the smallest part of the request. Everything before it is the cost.
Sources
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026: the two habits, the two usual causes of high spend,
/clearcosting nothing, the long-session warning, the/usageflags, the cache lifetimes. - Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026: the full context sent with every message, and each model having its own cache.
- Anthropic, Explore the context window (code.claude.com), read 4 October 2026: the simulated session and its token counts.
- Anthropic, Model configuration (code.claude.com), read 4 October 2026: Opus 5.5 as the default model.
- Anthropic, Pricing (platform.claude.com), read 4 October 2026: the token estimate and every list price in this post.


