How to reduce Claude Code token usage
To reduce Claude Code token usage, make the conversation shorter and keep repeated text at the cached price. Claude Code sends the whole conversation with every request, so usage grows with the length of the session, by Anthropic's documentation read on 4 October 2026. The changes, in the order we recommend: measure with /usage and /context, run /clear between unrelated tasks, keep CLAUDE.md under 200 lines, filter test and log output, use Sonnet for most work, and write requests that name the file. Anthropic states an average cost of $13 per developer per active day. This guide explains each change and links to a page on each.
Published October 4, 2026. Editorial.
Key takeaways
- Claude Code sends the full conversation with every request, so token usage grows with the length of the session, by Anthropic's documentation read on 4 October 2026.
- Anthropic states an average of $13 per developer per active day and $150 to $250 per developer per month across enterprise deployments (large organisations), with 90 percent of users below $30 per active day.
- Running `/clear` between unrelated tasks costs nothing, and Anthropic names long sessions that were never cleared as a usual cause of unexpectedly high spending.
- The default model on Anthropic's plans is Opus 5.5, and Anthropic advises Sonnet for most coding tasks because it costs less than Opus.
- In Anthropic's simulation of a session, a 45-token prompt leads to 6,900 tokens of file reads, so a request that names the file reduces more than a shorter request does.
Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. Every part of that work is counted in tokens, and the count grows with every step of a task. This guide explains why, and then gives the changes that reduce it, in the order we recommend making them.
The short version is this. Most of the tokens you pay for are text that the model reads again on every request, and you typed almost none of it. The largest savings therefore come from making the conversation shorter and from keeping repeated text at the lower cached price. Settings come second.
What a token is, and why usage grows with the conversation
A token is a piece of text that the model processes. Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters, or 0.75 of an English word [4]. File contents, command output, your requests and the model's replies are all counted this way.
The model keeps nothing between two requests. Anthropic's documentation, read on 4 October 2026, says that Claude Code therefore sends the full context each time: "the system prompt, your project context, every prior message and tool result, and your new message" [2]. The place that holds all of this is the context window, which is all the text the model reads in one request.
One request from you is often several requests to the model. Each time Claude reads a file or runs a command, Claude Code sends another request that contains the result together with the whole conversation [1]. A task that takes ten steps sends the conversation ten times, and it is longer each time. That count of ten is our invented example. The rule behind it is Anthropic's: "Token costs scale with context size: the more context Claude processes, the more tokens you use." [1]
Anthropic's simulation of a session shows the proportions. Seven items load before the user types anything, and they add up to 7,850 tokens by our arithmetic on Anthropic's figures. The user's prompt is 45 tokens. The four files Claude reads to answer it add 6,900 tokens [3]. Anthropic calls these "representative token counts", so they are an illustration, and your own session will differ.
Two mechanisms lower this cost, and Claude Code runs both without being asked [1]. The first is the prompt cache, a store of request text that the service has already processed. Text read again from the cache is billed at a lower price. On Claude Sonnet 5.5, for example, a cache read is listed at $0.20 per million tokens against $2 for new input, on Anthropic's pricing page read on 4 October 2026 [4]. The second is compaction, which replaces a long conversation history with a summary when the context window is close to full.
Ordinary habits can stop both mechanisms from helping. A session left open across unrelated tasks includes the text of all of them in every request. A change of model in the middle of a task starts a new cache, so the next request is processed in full. The rest of this guide is about those habits.
What Anthropic says Claude Code costs
Anthropic publishes three averages for Claude Code, all from its documentation read on 4 October 2026 [1]:
- Anthropic states an average of $13 per developer per active day across enterprise deployments, which are large organisations that use Claude Code.
- The same sentence gives $150 to $250 per developer per month.
- Costs stay below $30 per active day for 90 percent of users.
These are Anthropic's own figures about its own customers. No outside party has checked them, and Anthropic says that cost for each developer varies with the model chosen, the size of the code, and patterns of use such as running several sessions at once [1]. Use them as a first comparison. For a team budget, Anthropic's advice is to start with a small test group and record a baseline, which is a measurement taken before any change, before giving the tool to the whole team [1].
The figures matter in two different ways, depending on how you pay. If you pay per token, through an API key (a pay-per-token account) or a cloud provider, fewer tokens is a lower invoice. If you have a Pro, Max, Team or Enterprise plan, usage counts against the plan's limits, and the documentation warns that a session open for hours "can use far more of your plan limits than your activity suggests" [1]. In both cases the same changes apply.
Every change in one table
The table lists each change in this guide, what it reduces, and the page that explains it. The order is the order we recommend. Four names in it are explained in the sections below: CLAUDE.md (your instruction file), an MCP server (a program that connects Claude Code to an outside tool), a hook (a script that Claude Code runs by itself) and the effort level (how much the model reasons).
| Change | What it reduces | Page |
|---|---|---|
| Learn what a request contains | Guessing about where tokens go | Where Claude Code tokens go |
| Measure with two commands | Changes made without a baseline | How to read /usage and /context |
| Clear, compact or rewind at the moment that fits each command | The conversation history sent with every request | /clear, /compact or /rewind |
| Stop requests you did not type | Usage while you are away from the session | Why usage keeps rising in a long session |
| Shorten CLAUDE.md | Text loaded at the start of every session | How long CLAUDE.md should be |
| Prefer command-line tools to MCP servers | Tool listings sent with every request | MCP servers or command-line tools |
| Filter test and log output | Command output that enters the conversation | Hooks that trim test and log output |
| Match the model and effort level to the task | The price of each token, and the amount of reasoning | Which model and effort level to use |
| Write specific requests | File reads and repeat attempts | Writing requests that read fewer files |
| Keep a printed list | Forgetting the habits above | The Claude Code token checklist |
Start by measuring
Change nothing until you have two numbers. The command /usage shows what the session has used so far: the token counts for each model and an estimated cost. The command /context shows what is in the context window at this moment, by category [1][3]. The first tells you how much, and the second tells you where.
Read where Claude Code tokens go first if the idea of a request that contains the whole conversation is new to you. Then how to read /usage and /context explains each line of both screens. One caution from the documentation: the dollar figure in /usage is an estimate that Claude Code computes on your computer at list price, and the Usage page in the Claude Console, Anthropic's website for pay-per-token accounts, is the billing record [1].
Keep the conversation short
Anthropic names clearing between unrelated tasks as one of the two habits with the highest impact, and the /clear command itself costs nothing [1]. The same documentation names "long sessions that were never cleared" as one of the two usual causes of unexpectedly high spending on an API or cloud provider plan [1].
Three commands shorten a conversation, and each fits a different moment. /clear starts a new conversation with an empty context, for when the next task is unrelated. /compact replaces the history with a summary, for a pause inside one long task. /rewind returns to an earlier prompt, for when the last steps were a mistake. The page on /clear, /compact or /rewind compares what each one costs and what each one keeps.
A long session also grows in ways you do not see. The documentation lists eight causes, and several of them are requests that the session sends while you are away, such as a scheduled task that runs on its interval and sends the full context each time [1]. The page on why usage keeps rising in a long session shows how to find and stop each one.
Reduce what loads before you type
Some text is in every request from the first one. The largest part that you control is CLAUDE.md, the instruction file that Claude reads at the start of every session. Anthropic's target is under 200 lines for each file [5]. Instructions for one procedure belong in a skill, which is a packaged set of instructions that loads only when it is used [1]. The page on how long CLAUDE.md should be shows where to put each type of instruction.
The second part is outside tools. An MCP server is a program that connects Claude Code to an outside tool or data source, and each one adds its tool names and its instructions to the session [1]. Anthropic's advice is to prefer command-line tools such as gh and aws where they exist, because they add no listing of tools [1]. The page on MCP servers or command-line tools compares the two.
In Anthropic's simulation the start-up text is 7,850 tokens and stays that size, while the conversation grows with every file read and every command [3]. That is why this section comes after the one on conversation length. A short CLAUDE.md helps on every request, and a cleared conversation helps more.
Shorten what tools return
Command output enters the conversation in full, even when your screen shows one line. In Anthropic's simulation, one run of the test suite adds 1,200 tokens while the screen shows a short status line [3]. A hook is a script that Claude Code runs by itself at a fixed point, and a hook can filter a command's output before the model reads it. Anthropic's documentation says a hook that returns only the lines containing ERROR from a 10,000-line log reduces the context "from tens of thousands of tokens to hundreds" [1]. The page on hooks that trim test and log output reproduces Anthropic's example and lists what a filter can hide.
A subagent does similar work in a different way. A subagent is a second copy of Claude that works on one task in its own separate context window and sends back a summary. In Anthropic's simulation, a subagent reads 6,100 tokens of files and returns 420 tokens to the main conversation [3]. Its own requests still count in your usage [1].
Match the model and the effort level to the task
The model sets the price of every token. Anthropic's documentation says Sonnet handles most coding tasks well and costs less than Opus, and advises keeping Opus for complex architecture decisions or reasoning across many steps [1]. The default model on Anthropic's plans is Opus 5.5 [6], so this saving often starts with one command, /model sonnet.
The effort level is the setting for how much the model reasons before it replies. Reasoning is billed as output, the token type with the highest list price [1][4]. The page on which model and effort level to use gives the list prices and a table of tasks.
Choose both at the start of a session. Each model has its own prompt cache, so a switch in the middle of a task makes the next request read the whole conversation with no cache reads [2].
Write requests that need less reading
The wording of a request decides how much Claude has to search. Anthropic's documentation contrasts "improve this codebase", which causes broad scanning, with "add input validation to the login function in auth.ts", which needs few file reads [1]. Naming the file removes the search. Naming a check, such as a test that must pass, lets Claude find its own faults in the same turn, which removes repeat attempts.
For larger work, plan mode helps. Plan mode is a setting in which Claude reads the code and proposes an approach, and edits nothing until you approve. A wrong approach found in a plan costs one correction. The page on writing requests that read fewer files covers this, with examples of vague and specific requests.
The same reasoning applies to a whole project. Reveneau is an AI software development consultancy, all of its code is written by AI, and every change must pass an eval suite, which is a set of automated tests written from the specification, before release. The idea is the same at both sizes: a target that can be checked removes repeat attempts. The guide on eval-driven development explains that method.
Keep the prompt cache in use
Shortening the conversation reduces how many tokens are sent. The cache decides what each of those tokens costs. The two work together, and some actions help one and harm the other. /compact shortens the history and also builds a new cache entry. A model switch changes the price per token and starts a new cache.
Anthropic's rule for the cache is strict: it matches from the start of the request, and a change anywhere in that start means everything after it is processed again [2]. The cache also expires. It lasts one hour on a subscription, and five minutes on usage credits, an API key or a cloud provider by default [1]. The related guide on prompt caching in Claude Code explains which actions keep the cache and which discard it. This guide says what to do, and that guide explains how the cache works.
Limits and reports for a team
How an organisation limits spending depends on how it pays, by Anthropic's documentation read on 4 October 2026 [1]. On the Claude Console, an organisation sets workspace spend limits, and a dashboard shows spend for each member. On a Team or Enterprise plan, each member has a usage allowance. An administrator who turns on usage credits, which let a member keep working past that allowance, sets spend limits for the organisation, a group or one member, and a spend report shows estimated spend for each user. On Amazon Bedrock, Google Cloud's Agent Platform or Microsoft Foundry, the limits are in the cloud provider's own billing controls.
On a Pro or Max subscription, usage is included in the plan. Anthropic's documentation says the dollar figure in the Session block of /usage is intended for API users and is not relevant for a subscriber's billing [1].
One price question concerns large context windows. Some models can read 1 million tokens in one request. Anthropic's documentation says this window uses standard model pricing, with no extra charge for tokens beyond 200,000 [6]. Its pricing page gives the example that a 900,000-token request is billed at the same rate per token as a 9,000-token request [4]. A larger window still means larger requests if the conversation is never cleared.
What to do first, in ten minutes
These five steps need no installation and no administrator. Do them in this order.
- Run
/usageand/contextin the session you have open. Write down the token counts and the largest category. This is your baseline. - Check the model. If the session is on Opus and the work is everyday coding, run
/model sonnetat the start of your next session. - Run
/mcp. Disable each server you are not using. - Open CLAUDE.md and compare its length with the 200-line target. If it is longer, mark the step-by-step procedures. Those are the parts to move into skills later.
- Adopt one habit:
/rename, then/clear, each time you change to an unrelated task./resumereturns you to the old conversation if you need it [1].
After a week, run /usage again and compare. Then print the Claude Code token checklist, which puts every action in this guide on one list, grouped by when you do it.
What this guide covers
This guide covers Claude Code as Anthropic documents it on 4 October 2026. Every figure, command, setting name and version number comes from Anthropic's own documentation and pricing page, and each page lists its sources with the date they were read. Anthropic's figures are its own statements about its own product. Reveneau is independent of Anthropic.
This guide gives no figure for how much you will save, because the documentation gives none and the saving depends on your code, your model and your habits. Measure your own baseline, change one thing, and measure again.
Token use is a running cost of every Reveneau build, because Reveneau uses AI instead of hiring more engineers. For how that cost fits into the price of a whole project, see the guide on software development cost.
Explore the guide
Start here
Where Claude Code tokens go: what is sent on every message
Claude Code tokens go mostly to text the model reads again on every request. Each message re-sends the whole conversation: Anthropic's instructions, your CLAUDE.md instruction files, every earlier message, and every file and command result Claude has collected. In Anthropic's own simulation, read on 4 October 2026, the seven items that load before the user types add up to 7,850 tokens, and four file reads then add 6,900 tokens to a 45-token prompt. Input, output and cached tokens are billed at different prices, and the model's reasoning is billed as output. This page lists what is in the context window, when each part loads, and how to make each part smaller.
How to read /usage and /context in Claude Code
/usage and /context are the two measuring commands in Claude Code. /usage shows what the session has used so far: token counts for each model, an estimated cost that Claude Code computes on your computer at list price, and, on a paid plan, a breakdown of what counts against your limits. /context shows what is in the context window at this moment, by category. Anthropic's documentation, read on 4 October 2026, says the cost figure is an estimate and names the Usage page in the Claude Console as the billing record. This page explains each line, adds the status line for a continuous reading, and gives a five-minute routine.
The session
/clear, /compact or /rewind: which to use and what each costs
Use /clear when the next task is unrelated, /compact at a pause inside one long task, and /rewind when the last steps were a mistake. In Claude Code, /clear starts a new conversation and costs nothing, by Anthropic's documentation read on 4 October 2026. /compact replaces the history with a summary, and writing that summary is one request that reads the whole conversation. /rewind returns the conversation to an earlier prompt, and the next request reads text the prompt cache already holds. This page compares the three, explains automatic compaction and the auto-compact window, and lists what survives a compaction.
Why Claude Code usage keeps rising in a long session
Claude Code usage keeps rising in a long session because every request contains the whole conversation, and because the session can send requests while you are away. Anthropic's documentation, read on 4 October 2026, lists eight causes: long context, cache misses, scheduled tasks, messages from other sessions, goal check-ins, subagents and workflows, agent teammates, and compaction. On a paid plan, the /usage command flags a behaviour that accounts for 10 percent or more of recent usage. This page explains each cause in plain words, shows how to spot it, and names the setting or habit that stops it. Anthropic names running /clear between unrelated tasks as one of the two habits with the highest impact.
Setup
How long should CLAUDE.md be? What to keep and what to move into skills
A CLAUDE.md file should be under 200 lines. That is the target in Anthropic's documentation for Claude Code, read on 4 October 2026, and the reason is cost: CLAUDE.md loads at the start of every session and is then sent with every request, so each line is counted on every request, including requests that have no use for it. Keep the facts that every session needs, such as build commands and conventions. Move step-by-step procedures into skills, which load when they are used, and move instructions for one folder into path rules, which load when Claude opens a matching file. This page shows where to put each type of instruction.
MCP servers or command-line tools: what each adds to every message
An MCP server, a program that connects Claude Code to an outside service, adds its tool names and its instructions to every message. A command-line tool adds nothing until Claude runs it. That is the default behaviour in Anthropic's documentation, read on 4 October 2026: a feature named tool search holds back the full definition of each MCP tool until Claude needs it. When tool search is off, every definition loads at the start of the session and is sent with every request. Anthropic advises command-line tools such as `gh` and `aws` where they exist, because they add no tool listing. This page compares the two and shows how to check what your servers add.
Hooks that trim test and log output before Claude reads it
A hook can remove the passing lines from test output, and the lines without errors from a log, before Claude reads them. A hook is a command that Claude Code runs by itself at a fixed point, and a hook on the PreToolUse event can rewrite a terminal command before it runs. Anthropic's documentation, read on 4 October 2026, gives a working example that keeps only the failing lines of a test run. This page reproduces that example, shows how to check it with `/hooks` and a debug log, and lists what the filter can hide, such as a failure that is reported with an unexpected word. It also covers code intelligence plugins and subagents.
Choices
Which Claude Code model and effort level to use for which task
Use Sonnet at medium effort for most Claude Code work, Opus for architecture decisions, and Haiku for simple tasks given to a subagent. The choice of model is Anthropic's own advice, read on 4 October 2026, medium is the default effort on Sonnet 5.5, and the price list supports both: Claude Opus 5.5 is listed at $4 per million input tokens and $20 per million output tokens, twice the price of Claude Sonnet 5.5. The effort level sets how much the model reasons before it replies, and reasoning is billed as output. Opus 5.5 is the default model on Anthropic's plans, so the saving starts with one command, /model sonnet. This page gives the prices, the five effort levels and a table of tasks.
How to write a Claude Code request that reads fewer files
To write a Claude Code request that reads fewer files, name the file, the function, the change you want, and the check that shows the work is done. Anthropic's documentation, read on 4 October 2026, says a vague request such as "improve this codebase" causes broad scanning, while a request that names a function and a file needs few file reads. In Anthropic's simulation of a session, a 45-token prompt that names no file leads to four file reads of 6,900 tokens in total. This page gives examples of vague and specific requests, and covers plan mode, stopping early with Escape and /rewind, test targets, and giving exploration to a subagent.
Common questions
How do I reduce Claude Code token usage?
Reduce Claude Code token usage by keeping each conversation short and by keeping the prompt cache in use. Anthropic's documentation, read on 4 October 2026, says Claude Code sends the full conversation with every request, so the main habits are running `/clear` between unrelated tasks and choosing the model before the first request. After that, keep CLAUDE.md under 200 lines, filter long command output, and write requests that name the file.
How much does Claude Code cost per developer?
Claude Code costs an average of $13 per developer per active day across enterprise deployments (large organisations), by Anthropic's own figure, read on 4 October 2026. The same documentation gives $150 to $250 per developer per month, and says costs stay below $30 per active day for 90 percent of users. These are Anthropic's statements about its own customers. Your cost depends on the model, the size of the code and how you work.
Why does Claude Code use so many tokens?
Claude Code uses many tokens because most of each request is text the model reads again: the whole conversation is sent with every request. Anthropic's simulation of a session, read on 4 October 2026, shows the proportions. Seven items that add up to 7,850 tokens load before the user types, the user's prompt is 45 tokens, and the four files Claude reads to answer it add 6,900 tokens. Each later request contains all of that text again.
What should I change first to lower Claude Code usage?
Change your habit between tasks first: run `/rename` and then `/clear` each time you start unrelated work in Claude Code. Anthropic's documentation, read on 4 October 2026, says `/clear` costs nothing and that old context is sent again with every later message. The second change is the model. The default on Anthropic's plans is Opus 5.5, and Anthropic advises Sonnet for most coding tasks.
Does prompt caching reduce Claude Code costs?
Yes. Prompt caching reduces Claude Code costs, and Claude Code uses it without any setup. The prompt cache is a store of request text the service has already processed, and text read from it is billed at a lower price. On Anthropic's pricing page, read on 4 October 2026, a cache read on Claude Sonnet 5.5 is listed at $0.20 per million tokens against $2 for new input. Switching models in the middle of a task starts a new cache.
Do long sessions use up a Claude subscription limit faster?
Yes. A long Claude Code session uses a subscription's limits faster than a short one doing the same work. Anthropic's documentation, read on 4 October 2026, says a session that has been open for hours can use far more of your plan limits than your activity suggests, because a one-line question still counts usage for the whole conversation. On a paid plan, `/usage` flags long context when it accounts for 10 percent or more of recent usage.
Does Claude Code charge per token on a Pro or Max subscription?
On a Pro or Max subscription, Claude Code usage is included in the plan, and it counts against the plan's usage limits. Anthropic's documentation, read on 4 October 2026, says the dollar figure in the Session block of `/usage` is intended for API users and has no bearing on a subscriber's billing. Usage credits, if you turn them on, let you keep working past the plan's limit.
How do I estimate Claude Code costs for a team?
Estimate Claude Code costs for a team by measuring a small test group first. Anthropic's documentation, read on 4 October 2026, advises starting with a small group and using its tracking tools to measure usage before giving the tool to everyone. Its published average of $13 per developer per active day is a first comparison. On a Team or Enterprise plan the spend report shows estimated spend for each user, and on the Claude Console the dashboard shows spend for each member.
Does a smaller CLAUDE.md lower Claude Code costs?
Yes. A smaller CLAUDE.md lowers Claude Code costs, because the file loads at the start of every session and is then part of every request. Anthropic's documentation, read on 4 October 2026, sets a target of under 200 lines for each file and advises moving instructions for specific procedures into skills, which load only when used. The effect is smaller than clearing the conversation, since the conversation has no fixed size.
Can an administrator limit Claude Code spending?
Yes. An administrator can limit Claude Code spending in the place that matches how the organisation pays. Anthropic's documentation, read on 4 October 2026, says organisations on the Claude Console, its website for pay-per-token accounts, set workspace spend limits, and Team and Enterprise plans set spend limits on usage credits for the organisation, a group or one member. On Amazon Bedrock, Google Cloud or Microsoft Foundry, the limits are in the cloud provider's own billing controls.
Does the 1 million token context window cost more per token?
The 1 million token context window in Claude Code is billed at the standard price per token. Anthropic's documentation, read on 4 October 2026, says it uses standard model pricing with no extra charge for tokens beyond 200,000, and the pricing page gives the example that a 900,000-token request is billed at the same rate per token as a 9,000-token request. A larger window still means larger requests if the conversation is never cleared.
References
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Explore the context window (code.claude.com), read 4 October 2026
- Anthropic, Pricing (platform.claude.com), read 4 October 2026
- Anthropic, How Claude remembers your project (code.claude.com), read 4 October 2026
- Anthropic, Model configuration (code.claude.com), read 4 October 2026