How to reduce Claude Code token usage / The session
Why Claude Code usage keeps rising in a long session
Claude Code usage keeps rising in a long session because every request contains the whole conversation, and because the session can send requests while you are away. Anthropic's documentation, read on 4 October 2026, lists eight causes: long context, cache misses, scheduled tasks, messages from other sessions, goal check-ins, subagents and workflows, agent teammates, and compaction. On a paid plan, the /usage command flags a behaviour that accounts for 10 percent or more of recent usage. This page explains each cause in plain words, shows how to spot it, and names the setting or habit that stops it. Anthropic names running /clear between unrelated tasks as one of the two habits with the highest impact.
Published October 4, 2026. Editorial.
Key takeaways
- Anthropic's documentation, read on 4 October 2026, lists eight causes of rising usage in a long Claude Code session, and five of them are requests that you did not type.
- By default the prompt cache lasts one hour on a subscription and five minutes on usage credits, a pay-per-token key or a cloud provider, and the first message after a longer break reprocesses the full context.
- A scheduled task in Claude Code sends the full context each time it runs, even in an idle session, and a repeating task expires seven days after it was created.
- From Claude Code v2.1.246, a goal starts at most three idle check-ins between your prompts, and setting `CLAUDE_CODE_GOAL_CHECKIN_MINUTES` to `0` turns check-ins off.
- Background usage in Claude Code is typically under $0.04 per session by Anthropic's figure, so reducing it saves little.
A session in Claude Code, Anthropic's coding tool, can stay open for a whole working day. Anthropic's documentation, read on 4 October 2026, warns that such a session "can use far more of your plan limits than your activity suggests", and it lists eight reasons [1]. This page explains each one in plain words, shows how to see it, and names the setting or habit that stops it.
Four terms are needed first. A token is a piece of text that a model processes, and all usage is counted in tokens. A request is one message sent to the model. The context window is the largest amount of text the model can read in one request. The prompt cache is a store, kept by the service that runs the model, of request text it has already processed, so that unchanged text can be read again at a lower price.
The eight causes belong to two groups. Three are about the size of the conversation: long context, cache misses and compaction. Five are about requests that Claude Code sends when you did not type anything: scheduled tasks, messages from other sessions, goal check-ins, subagents and teammates. The second group is harder to notice, because the session uses tokens while you are away. This page belongs to the guide on how to reduce Claude Code token usage.
The eight causes: how to spot each one and how to fix it
| Cause | How to spot it | Fix |
|---|---|---|
| 1. Long context | A high percentage in /context, or a "long context" flag in /usage [1] |
/clear between unrelated tasks |
| 2. Cache misses | Misses on the Prompt cache (main) line of /usage, or a "cache misses" flag [1] |
Start a new conversation after a long break, or lengthen the cache lifetime |
| 3. Scheduled tasks | The Loops rows in /usage [1] |
Cancel the task, or press Esc on a waiting loop [2] |
| 4. Messages from other sessions | A new turn that starts while this session is idle [1] | Set crossSessionInbound to hold [1] |
| 5. Goal check-ins | /goal shows an active goal [3] |
/goal clear, or set CLAUDE_CODE_GOAL_CHECKIN_MINUTES to 0 [1] |
| 6. Subagents and workflows | The subagent share in the /usage attribution [1] |
Give simple subagent tasks a cheaper model [1] |
| 7. Agent teammates | The agent team feature is on and teammates are still running | Shut down each teammate when its work is done [1] |
| 8. Compaction | An expected rebuild on the Prompt cache (main) line of /usage [1] |
/clear when you do not need the old conversation [1] |
The page on how to read /usage and /context explains the screens named in the second column. On a Pro, Max, Team or Enterprise plan, /usage flags a behaviour "such as long context or cache misses" when it accounts for 10 percent or more of your recent usage, and gives a tip with each flag [1]. Start there.
1. Long context: every request contains the whole conversation
The documentation states the mechanism in one sentence: "Claude Code sends your full conversation with every request, and each time Claude uses tools it sends another request carrying that batch of tool results." [1] A tool is an action such as reading a file or running a command. One message from you can therefore produce many requests, and each one contains everything said so far.
The prompt cache lowers the price of this repeated reading and leaves the count unchanged. Anthropic's example is a one-line question in a session that has been open all day: it is still counted as usage for the whole conversation [1].
Here is an invented example, to show the size of the effect. Suppose a conversation has grown to 400,000 tokens on Claude Sonnet 4.6. Anthropic's pricing page, read on 4 October 2026, lists cache reads for that model at $0.30 per million tokens [7]. One request then costs $0.12 for the repeated reading alone. A task in which Claude uses tools 20 times sends 20 requests, which is $2.40 before any new work is counted.
The fix is to keep the conversation short. Run /clear when you change to unrelated work. The page on /clear, /compact and /rewind compares the commands that do this.
2. Cache misses: the first message after a break
Stored text in the prompt cache expires when it is unused for a set time. Anthropic's documentation, read on 4 October 2026, gives the times: one hour on a subscription, and five minutes once you are using usage credits, which are paid usage beyond your plan's limit. With a pay-per-token key, or through a cloud provider such as Amazon Bedrock, the default is five minutes [1].
When the stored text has expired, the next request is a cache miss. The documentation says that your first message after such a break "misses the cache and reprocesses your full context" [1]. Continue the invented example. The same 400,000 tokens, written to the five-minute cache at the listed $3.75 per million tokens, cost $1.50 for one request [7]. That is 12.5 times the $0.12 of a request that reads from the cache.
A lunch break on a five-minute cache is one miss. A day with ten breaks is ten.
The fixes, by Anthropic's documentation:
- After a long break, begin a new conversation with
/clearif the old one is no longer needed. A new conversation has little to process. - On a Pro or Max plan, when you resume a session that is over 100,000 tokens and has been inactive for a period Anthropic gives as "more than about an hour", Claude Code offers to resume from a summary, so that later requests contain the summary in place of the full history [4].
- With a pay-per-token key or a cloud provider, the setting
promptCacheTtlwith the value1hgives the main conversation a one-hour cache. It requires Claude Code v2.1.242 or later. The one-hour cache has a higher price for each write, so it suits work with breaks and costs more for work without them [6].
A break is one of several causes of a miss. Changing the model in the middle of a task is another. The page on what invalidates the prompt cache lists them all, and cache lifetime: five minutes or one hour covers the choice between the two times.
3. Scheduled tasks: a prompt that repeats while you are away
A scheduled task is a prompt that Claude runs again on an interval. The usual way to start one is the /loop command, for example /loop 5m check the deploy, which asks Claude every five minutes whether a release has finished [2].
The cost is in one line of the documentation: a scheduled task "fires on its interval even while the session is idle, sending your full context each time" [1]. An invented example: a five-minute loop left running for eight hours can run up to 96 times, which is 8 hours times 12 runs an hour. Each run sends the whole conversation.
Three limits from the scheduled-tasks documentation, read on 4 October 2026, limit how far this can go [2]:
- The shortest interval for
/loopis one minute. - A session can have up to 50 scheduled tasks at the same time.
- A repeating task expires seven days after it was created.
How to see it. The /usage breakdown has a row for each of the scheduled tasks that used the most tokens, with its total tokens and tokens per run. This requires Claude Code v2.1.242 or later [1]. You can also ask Claude in plain words: "what scheduled tasks do I have?" [2]
The fixes. Ask Claude to cancel the task by name, for example "cancel the deploy check job". For a loop where Claude chooses the interval itself, press Esc while it is waiting. To turn the feature off completely, set CLAUDE_CODE_DISABLE_CRON=1 [2].
Start every loop in a short conversation. A loop that runs inside a long conversation sends that whole conversation on every run. Where it is available, the documentation points to a cheaper tool for waiting on an event: the Monitor tool runs a script in the background and passes each line of its output back, which Anthropic says "is often more token-efficient" than running a prompt again on an interval [2].
4. Messages from your other sessions
You can have several Claude Code sessions open, and one can send a message to another. The documentation says that Claude Code delivers such a message "as a new turn when this session sits idle, sending your full context each time" [1]. A long, idle session that receives five messages sends its whole conversation five times.
The fix is the setting crossSessionInbound with the value hold, which holds incoming messages and stops them from starting a turn [1]. Anthropic's cost documentation gives no screen that counts these messages separately.
5. Goal check-ins
The /goal command sets a condition, such as "all tests pass", and Claude keeps working turn after turn until the condition is met. After each turn, a small, fast model checks whether the condition holds [3]. Anthropic describes the tokens for that check as "typically negligible compared to main-turn spend" [3].
The larger cost is the check-in. A goal can be left waiting on work that runs in the background. By Anthropic's documentation, read on 4 October 2026, a check-in is due once that work has kept the goal waiting for 30 minutes. Claude Code then asks Claude to look at the running work. In an idle session it starts that turn by itself, and the turn sends the full context [1][3]. Later check-ins come 1 hour after the first and then every 2 hours [3].
Three version facts from the same documentation [1][3]:
- Idle check-ins require Claude Code v2.1.236 or later.
- From v2.1.246, Claude Code starts at most three idle check-ins for each goal between your prompts. Before that version there was no upper limit.
- Setting
CLAUDE_CODE_GOAL_CHECKIN_MINUTESto0turns check-ins off.
The fix for a goal you no longer need is /goal clear. Starting a new conversation with /clear also removes an active goal [3].
6. Subagents and workflows
A subagent is a second copy of Claude that does one task in its own context window and returns a summary. A workflow is a script, written by Claude, that starts many subagents. The documentation says that every subagent, and every agent that a workflow starts, "sends its own requests on top of the main conversation's" [1].
Subagents have a real benefit: the files they read are kept outside your main conversation. Their own usage is still yours. A subagent also begins without the main conversation's stored text, and its cache lifetime is five minutes even on a subscription [6].
How to see it. The attribution part of the /usage breakdown shows the share that subagents used [1].
The fix is to match the model to the job. Anthropic's advice for simple subagent tasks is to set model: haiku in the subagent's configuration, which selects Haiku, the model family with the lowest list prices on Anthropic's pricing page [1][7]. The page on which model and effort level to use covers this choice.
7. Agent teammates
An agent team is a group of Claude Code instances that work together, each with its own context window. The feature is off by default and is turned on with CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 [1].
The documentation states the cost plainly: "Each active teammate continues consuming tokens until it exits or the session ends." [1] Anthropic puts the usage of an agent team at 7 times that of a standard session when the teammates run in plan mode, which is the mode where Claude studies the code and proposes a plan before it edits [1].
The fixes from the same page: keep teams small, use Sonnet for teammates, keep the starting instructions short, and shut down each teammate when its work is done [1].
8. Compaction
Compaction replaces the conversation history with a summary. It reduces cause 1 and it has its own cost. The documentation explains: "/compact reads the conversation it summarizes, so compacting a large context is itself a large request." It then adds: "When you want a fresh start instead of continuity, /clear costs nothing" [1].
A compaction is one large request that makes every later request smaller. It is the right choice when the same task continues. When the next task is new, /clear gives the same smaller requests for no cost. The page on what /compact costs gives the detail.
Background usage and prompt suggestions are small
Two further sources of usage exist, and both are small by Anthropic's own figures.
Background usage. Claude Code uses some tokens when idle: it summarises earlier conversations so that claude --resume can list them, and some commands, such as /usage, send a request to check status. Anthropic states that this is typically under $0.04 per session [1].
Prompt suggestions. After Claude responds, Claude Code can show a suggested next prompt. Each suggestion is a short request to the same model your session uses. Anthropic's documentation says it reuses the conversation's stored text, "so it is mostly cache reads plus a few output tokens" [1]. Claude Code skips the suggestion when the stored text has expired, to avoid cost, and when your account is close to its limit [5]. To turn suggestions off, switch off Prompt suggestions in /config, or set promptSuggestionEnabled to false in your settings file [5].
Our position: leave prompt suggestions on until the eight causes above are handled. Background usage of under $0.04 a session is small next to the $1.50 of the single cache miss in the invented example above.
A routine for the end of the day
Reveneau is an AI software development consultancy, and all of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. This is what Reveneau recommends before you leave a session:
- Run
/usage, pressw, and read the flags and the Loops rows. - Ask Claude which scheduled tasks exist, and cancel the ones you do not need overnight.
- Run
/goalto check for an active goal, and clear it if the work is done. - Shut down agent teammates.
- Run
/clear, or close the session.
Anthropic's documentation names "long sessions that were never cleared" as one of the two usual sources of unexpectedly high spending on a pay-per-token plan [1]. Step 5 removes that source. The Claude Code token checklist has a printable list of actions for the whole guide.
Common questions
Why does Claude Code use tokens when I am not typing?
Claude Code uses tokens when you are not typing because several features start a turn by themselves, and each turn sends the full context. Anthropic's documentation, read on 4 October 2026, names scheduled tasks, messages from your other sessions, goal check-ins, subagents and agent teammates. A scheduled task, for example, runs on its interval even while the session is idle. Cancel loops and clear goals before you leave a session open.
Why is the first message after a break slow and expensive in Claude Code?
The first message after a break is slow and expensive because the stored text in the prompt cache has expired, so the whole conversation is processed again in full. Anthropic's documentation, read on 4 October 2026, says this happens after a break longer than the cache lifetime: one hour on a subscription, and five minutes on usage credits, a pay-per-token key or a cloud provider.
How long can I pause before Claude Code reprocesses the whole conversation?
You can pause for up to one hour on a subscription that is within its included usage, and for up to five minutes otherwise. Anthropic's documentation, read on 4 October 2026, gives five minutes as the default with a pay-per-token key or a cloud provider, and for a subscription that is using usage credits. The setting `promptCacheTtl` with the value `1h` lengthens the time, and requires Claude Code v2.1.242 or later.
Does /loop use tokens while the session is idle?
Yes. A task started with `/loop` runs on its interval even while the session is idle, and Anthropic's documentation, read on 4 October 2026, says it sends your full context each time. In an invented example, a five-minute loop left for eight hours can run up to 96 times. Start a loop in a short conversation, and cancel it when the thing it was waiting for has happened.
How do I see which scheduled tasks are using tokens?
You see scheduled tasks and their token use in the Loops rows of the `/usage` breakdown. Each row gives how often the task runs, how many times it ran, its total tokens and its tokens per run. Anthropic's documentation, read on 4 October 2026, says these rows require Claude Code v2.1.242 or later. You can also ask Claude in plain words which scheduled tasks you have.
How do I stop a /loop in Claude Code?
You stop a `/loop` by asking Claude to cancel the task, or by pressing `Esc` if it is a loop where Claude chooses the interval and it is waiting for its next run. Anthropic's documentation, read on 4 October 2026, says a loop on a fixed interval keeps running until you cancel it or seven days pass. Setting `CLAUDE_CODE_DISABLE_CRON=1` turns the scheduling feature off completely.
What is a goal check-in in Claude Code?
A goal check-in is a turn in which Claude Code asks Claude to look at background work that has kept a goal waiting. Anthropic's documentation, read on 4 October 2026, says the first check-in is due after 30 minutes, and that in an idle session Claude Code starts the turn by itself, which sends the full context. From v2.1.246 there are at most three idle check-ins for each goal between your prompts.
Do subagents count against my usage limit?
Yes. Every subagent sends its own requests, and those requests count against the same usage as your main conversation. Anthropic's documentation, read on 4 October 2026, says the attribution part of the `/usage` breakdown shows the share that subagents used. For simple subagent tasks, Anthropic advises setting `model: haiku` in the subagent's configuration so that the work runs on a cheaper model.
How much more do agent teams use than a normal session?
Anthropic puts the token use of an agent team at 7 times that of a standard session when the teammates run in plan mode. The reason given in Anthropic's documentation, read on 4 October 2026, is that each teammate is a separate copy of Claude with its own context window. Each active teammate keeps using tokens until it exits, so shut teammates down when their work is done.
Do prompt suggestions cost tokens?
Yes. Each prompt suggestion in Claude Code is a short request to the model your session is using, and it counts toward your plan limits or your pay-per-token costs. Anthropic's documentation, read on 4 October 2026, says the request reuses the conversation's prompt cache, so it is mostly cache reads plus a few output tokens. You can turn suggestions off in `/config` or with `promptSuggestionEnabled` set to `false`.
How much does Claude Code use in the background?
Background usage in Claude Code is typically under $0.04 per session, by Anthropic's own figure in documentation read on 4 October 2026. It covers two things: summaries of earlier conversations, which `claude --resume` uses to list them, and status requests sent by some commands such as `/usage`. This amount is small next to the cost of one long conversation, so look at the eight main causes first.
What should I do before leaving a Claude Code session open overnight?
Before you leave a Claude Code session open overnight, cancel scheduled tasks, clear any active goal, shut down agent teammates, and run `/clear`. Each of these stops a source of requests that Anthropic's documentation, read on 4 October 2026, says can run while the session is idle. A repeating scheduled task expires only after seven days, so a forgotten loop can run for a week.
References
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Run prompts on a schedule (code.claude.com), read 4 October 2026
- Anthropic, Keep Claude working toward a goal (code.claude.com), read 4 October 2026
- Anthropic, Manage sessions (code.claude.com), read 4 October 2026
- Anthropic, Interactive mode (code.claude.com), read 4 October 2026
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Pricing (platform.claude.com), read 4 October 2026