Guide

How to reduce Claude Code token usage

To reduce Claude Code token usage, make the conversation shorter and keep repeated text at the cached price. Claude Code sends the whole conversation with every request, so usage grows with the length of the session, by Anthropic's documentation read on 4 October 2026. The changes, in the order we recommend: measure with /usage and /context, run /clear between unrelated tasks, keep CLAUDE.md under 200 lines, filter test and log output, use Sonnet for most work, and write requests that name the file. Anthropic states an average cost of $13 per developer per active day. This guide explains each change and links to a page on each.

Published October 4, 2026. Editorial.

Key takeaways

  • Claude Code sends the full conversation with every request, so token usage grows with the length of the session, by Anthropic's documentation read on 4 October 2026.
  • Anthropic states an average of $13 per developer per active day and $150 to $250 per developer per month across enterprise deployments (large organisations), with 90 percent of users below $30 per active day.
  • Running `/clear` between unrelated tasks costs nothing, and Anthropic names long sessions that were never cleared as a usual cause of unexpectedly high spending.
  • The default model on Anthropic's plans is Opus 5.5, and Anthropic advises Sonnet for most coding tasks because it costs less than Opus.
  • In Anthropic's simulation of a session, a 45-token prompt leads to 6,900 tokens of file reads, so a request that names the file reduces more than a shorter request does.

Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. Every part of that work is counted in tokens, and the count grows with every step of a task. This guide explains why, and then gives the changes that reduce it, in the order we recommend making them.

The short version is this. Most of the tokens you pay for are text that the model reads again on every request, and you typed almost none of it. The largest savings therefore come from making the conversation shorter and from keeping repeated text at the lower cached price. Settings come second.

What a token is, and why usage grows with the conversation

A token is a piece of text that the model processes. Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters, or 0.75 of an English word [4]. File contents, command output, your requests and the model's replies are all counted this way.

The model keeps nothing between two requests. Anthropic's documentation, read on 4 October 2026, says that Claude Code therefore sends the full context each time: "the system prompt, your project context, every prior message and tool result, and your new message" [2]. The place that holds all of this is the context window, which is all the text the model reads in one request.

One request from you is often several requests to the model. Each time Claude reads a file or runs a command, Claude Code sends another request that contains the result together with the whole conversation [1]. A task that takes ten steps sends the conversation ten times, and it is longer each time. That count of ten is our invented example. The rule behind it is Anthropic's: "Token costs scale with context size: the more context Claude processes, the more tokens you use." [1]

Anthropic's simulation of a session shows the proportions. Seven items load before the user types anything, and they add up to 7,850 tokens by our arithmetic on Anthropic's figures. The user's prompt is 45 tokens. The four files Claude reads to answer it add 6,900 tokens [3]. Anthropic calls these "representative token counts", so they are an illustration, and your own session will differ.

Two mechanisms lower this cost, and Claude Code runs both without being asked [1]. The first is the prompt cache, a store of request text that the service has already processed. Text read again from the cache is billed at a lower price. On Claude Sonnet 5.5, for example, a cache read is listed at $0.20 per million tokens against $2 for new input, on Anthropic's pricing page read on 4 October 2026 [4]. The second is compaction, which replaces a long conversation history with a summary when the context window is close to full.

Ordinary habits can stop both mechanisms from helping. A session left open across unrelated tasks includes the text of all of them in every request. A change of model in the middle of a task starts a new cache, so the next request is processed in full. The rest of this guide is about those habits.

What Anthropic says Claude Code costs

Anthropic publishes three averages for Claude Code, all from its documentation read on 4 October 2026 [1]:

  • Anthropic states an average of $13 per developer per active day across enterprise deployments, which are large organisations that use Claude Code.
  • The same sentence gives $150 to $250 per developer per month.
  • Costs stay below $30 per active day for 90 percent of users.

These are Anthropic's own figures about its own customers. No outside party has checked them, and Anthropic says that cost for each developer varies with the model chosen, the size of the code, and patterns of use such as running several sessions at once [1]. Use them as a first comparison. For a team budget, Anthropic's advice is to start with a small test group and record a baseline, which is a measurement taken before any change, before giving the tool to the whole team [1].

The figures matter in two different ways, depending on how you pay. If you pay per token, through an API key (a pay-per-token account) or a cloud provider, fewer tokens is a lower invoice. If you have a Pro, Max, Team or Enterprise plan, usage counts against the plan's limits, and the documentation warns that a session open for hours "can use far more of your plan limits than your activity suggests" [1]. In both cases the same changes apply.

Every change in one table

The table lists each change in this guide, what it reduces, and the page that explains it. The order is the order we recommend. Four names in it are explained in the sections below: CLAUDE.md (your instruction file), an MCP server (a program that connects Claude Code to an outside tool), a hook (a script that Claude Code runs by itself) and the effort level (how much the model reasons).

Change What it reduces Page
Learn what a request contains Guessing about where tokens go Where Claude Code tokens go
Measure with two commands Changes made without a baseline How to read /usage and /context
Clear, compact or rewind at the moment that fits each command The conversation history sent with every request /clear, /compact or /rewind
Stop requests you did not type Usage while you are away from the session Why usage keeps rising in a long session
Shorten CLAUDE.md Text loaded at the start of every session How long CLAUDE.md should be
Prefer command-line tools to MCP servers Tool listings sent with every request MCP servers or command-line tools
Filter test and log output Command output that enters the conversation Hooks that trim test and log output
Match the model and effort level to the task The price of each token, and the amount of reasoning Which model and effort level to use
Write specific requests File reads and repeat attempts Writing requests that read fewer files
Keep a printed list Forgetting the habits above The Claude Code token checklist

Start by measuring

Change nothing until you have two numbers. The command /usage shows what the session has used so far: the token counts for each model and an estimated cost. The command /context shows what is in the context window at this moment, by category [1][3]. The first tells you how much, and the second tells you where.

Read where Claude Code tokens go first if the idea of a request that contains the whole conversation is new to you. Then how to read /usage and /context explains each line of both screens. One caution from the documentation: the dollar figure in /usage is an estimate that Claude Code computes on your computer at list price, and the Usage page in the Claude Console, Anthropic's website for pay-per-token accounts, is the billing record [1].

Keep the conversation short

Anthropic names clearing between unrelated tasks as one of the two habits with the highest impact, and the /clear command itself costs nothing [1]. The same documentation names "long sessions that were never cleared" as one of the two usual causes of unexpectedly high spending on an API or cloud provider plan [1].

Three commands shorten a conversation, and each fits a different moment. /clear starts a new conversation with an empty context, for when the next task is unrelated. /compact replaces the history with a summary, for a pause inside one long task. /rewind returns to an earlier prompt, for when the last steps were a mistake. The page on /clear, /compact or /rewind compares what each one costs and what each one keeps.

A long session also grows in ways you do not see. The documentation lists eight causes, and several of them are requests that the session sends while you are away, such as a scheduled task that runs on its interval and sends the full context each time [1]. The page on why usage keeps rising in a long session shows how to find and stop each one.

Reduce what loads before you type

Some text is in every request from the first one. The largest part that you control is CLAUDE.md, the instruction file that Claude reads at the start of every session. Anthropic's target is under 200 lines for each file [5]. Instructions for one procedure belong in a skill, which is a packaged set of instructions that loads only when it is used [1]. The page on how long CLAUDE.md should be shows where to put each type of instruction.

The second part is outside tools. An MCP server is a program that connects Claude Code to an outside tool or data source, and each one adds its tool names and its instructions to the session [1]. Anthropic's advice is to prefer command-line tools such as gh and aws where they exist, because they add no listing of tools [1]. The page on MCP servers or command-line tools compares the two.

In Anthropic's simulation the start-up text is 7,850 tokens and stays that size, while the conversation grows with every file read and every command [3]. That is why this section comes after the one on conversation length. A short CLAUDE.md helps on every request, and a cleared conversation helps more.

Shorten what tools return

Command output enters the conversation in full, even when your screen shows one line. In Anthropic's simulation, one run of the test suite adds 1,200 tokens while the screen shows a short status line [3]. A hook is a script that Claude Code runs by itself at a fixed point, and a hook can filter a command's output before the model reads it. Anthropic's documentation says a hook that returns only the lines containing ERROR from a 10,000-line log reduces the context "from tens of thousands of tokens to hundreds" [1]. The page on hooks that trim test and log output reproduces Anthropic's example and lists what a filter can hide.

A subagent does similar work in a different way. A subagent is a second copy of Claude that works on one task in its own separate context window and sends back a summary. In Anthropic's simulation, a subagent reads 6,100 tokens of files and returns 420 tokens to the main conversation [3]. Its own requests still count in your usage [1].

Match the model and the effort level to the task

The model sets the price of every token. Anthropic's documentation says Sonnet handles most coding tasks well and costs less than Opus, and advises keeping Opus for complex architecture decisions or reasoning across many steps [1]. The default model on Anthropic's plans is Opus 5.5 [6], so this saving often starts with one command, /model sonnet.

The effort level is the setting for how much the model reasons before it replies. Reasoning is billed as output, the token type with the highest list price [1][4]. The page on which model and effort level to use gives the list prices and a table of tasks.

Choose both at the start of a session. Each model has its own prompt cache, so a switch in the middle of a task makes the next request read the whole conversation with no cache reads [2].

Write requests that need less reading

The wording of a request decides how much Claude has to search. Anthropic's documentation contrasts "improve this codebase", which causes broad scanning, with "add input validation to the login function in auth.ts", which needs few file reads [1]. Naming the file removes the search. Naming a check, such as a test that must pass, lets Claude find its own faults in the same turn, which removes repeat attempts.

For larger work, plan mode helps. Plan mode is a setting in which Claude reads the code and proposes an approach, and edits nothing until you approve. A wrong approach found in a plan costs one correction. The page on writing requests that read fewer files covers this, with examples of vague and specific requests.

The same reasoning applies to a whole project. Reveneau is an AI software development consultancy, all of its code is written by AI, and every change must pass an eval suite, which is a set of automated tests written from the specification, before release. The idea is the same at both sizes: a target that can be checked removes repeat attempts. The guide on eval-driven development explains that method.

Keep the prompt cache in use

Shortening the conversation reduces how many tokens are sent. The cache decides what each of those tokens costs. The two work together, and some actions help one and harm the other. /compact shortens the history and also builds a new cache entry. A model switch changes the price per token and starts a new cache.

Anthropic's rule for the cache is strict: it matches from the start of the request, and a change anywhere in that start means everything after it is processed again [2]. The cache also expires. It lasts one hour on a subscription, and five minutes on usage credits, an API key or a cloud provider by default [1]. The related guide on prompt caching in Claude Code explains which actions keep the cache and which discard it. This guide says what to do, and that guide explains how the cache works.

Limits and reports for a team

How an organisation limits spending depends on how it pays, by Anthropic's documentation read on 4 October 2026 [1]. On the Claude Console, an organisation sets workspace spend limits, and a dashboard shows spend for each member. On a Team or Enterprise plan, each member has a usage allowance. An administrator who turns on usage credits, which let a member keep working past that allowance, sets spend limits for the organisation, a group or one member, and a spend report shows estimated spend for each user. On Amazon Bedrock, Google Cloud's Agent Platform or Microsoft Foundry, the limits are in the cloud provider's own billing controls.

On a Pro or Max subscription, usage is included in the plan. Anthropic's documentation says the dollar figure in the Session block of /usage is intended for API users and is not relevant for a subscriber's billing [1].

One price question concerns large context windows. Some models can read 1 million tokens in one request. Anthropic's documentation says this window uses standard model pricing, with no extra charge for tokens beyond 200,000 [6]. Its pricing page gives the example that a 900,000-token request is billed at the same rate per token as a 9,000-token request [4]. A larger window still means larger requests if the conversation is never cleared.

What to do first, in ten minutes

These five steps need no installation and no administrator. Do them in this order.

  1. Run /usage and /context in the session you have open. Write down the token counts and the largest category. This is your baseline.
  2. Check the model. If the session is on Opus and the work is everyday coding, run /model sonnet at the start of your next session.
  3. Run /mcp. Disable each server you are not using.
  4. Open CLAUDE.md and compare its length with the 200-line target. If it is longer, mark the step-by-step procedures. Those are the parts to move into skills later.
  5. Adopt one habit: /rename, then /clear, each time you change to an unrelated task. /resume returns you to the old conversation if you need it [1].

After a week, run /usage again and compare. Then print the Claude Code token checklist, which puts every action in this guide on one list, grouped by when you do it.

What this guide covers

This guide covers Claude Code as Anthropic documents it on 4 October 2026. Every figure, command, setting name and version number comes from Anthropic's own documentation and pricing page, and each page lists its sources with the date they were read. Anthropic's figures are its own statements about its own product. Reveneau is independent of Anthropic.

This guide gives no figure for how much you will save, because the documentation gives none and the saving depends on your code, your model and your habits. Measure your own baseline, change one thing, and measure again.

Token use is a running cost of every Reveneau build, because Reveneau uses AI instead of hiring more engineers. For how that cost fits into the price of a whole project, see the guide on software development cost.

Explore the guide

Setup

How long should CLAUDE.md be? What to keep and what to move into skills

A CLAUDE.md file should be under 200 lines. That is the target in Anthropic's documentation for Claude Code, read on 4 October 2026, and the reason is cost: CLAUDE.md loads at the start of every session and is then sent with every request, so each line is counted on every request, including requests that have no use for it. Keep the facts that every session needs, such as build commands and conventions. Move step-by-step procedures into skills, which load when they are used, and move instructions for one folder into path rules, which load when Claude opens a matching file. This page shows where to put each type of instruction.

MCP servers or command-line tools: what each adds to every message

An MCP server, a program that connects Claude Code to an outside service, adds its tool names and its instructions to every message. A command-line tool adds nothing until Claude runs it. That is the default behaviour in Anthropic's documentation, read on 4 October 2026: a feature named tool search holds back the full definition of each MCP tool until Claude needs it. When tool search is off, every definition loads at the start of the session and is sent with every request. Anthropic advises command-line tools such as `gh` and `aws` where they exist, because they add no tool listing. This page compares the two and shows how to check what your servers add.

Hooks that trim test and log output before Claude reads it

A hook can remove the passing lines from test output, and the lines without errors from a log, before Claude reads them. A hook is a command that Claude Code runs by itself at a fixed point, and a hook on the PreToolUse event can rewrite a terminal command before it runs. Anthropic's documentation, read on 4 October 2026, gives a working example that keeps only the failing lines of a test run. This page reproduces that example, shows how to check it with `/hooks` and a debug log, and lists what the filter can hide, such as a failure that is reported with an unexpected word. It also covers code intelligence plugins and subagents.

Common questions

How do I reduce Claude Code token usage?

Reduce Claude Code token usage by keeping each conversation short and by keeping the prompt cache in use. Anthropic's documentation, read on 4 October 2026, says Claude Code sends the full conversation with every request, so the main habits are running `/clear` between unrelated tasks and choosing the model before the first request. After that, keep CLAUDE.md under 200 lines, filter long command output, and write requests that name the file.

How much does Claude Code cost per developer?

Claude Code costs an average of $13 per developer per active day across enterprise deployments (large organisations), by Anthropic's own figure, read on 4 October 2026. The same documentation gives $150 to $250 per developer per month, and says costs stay below $30 per active day for 90 percent of users. These are Anthropic's statements about its own customers. Your cost depends on the model, the size of the code and how you work.

Why does Claude Code use so many tokens?

Claude Code uses many tokens because most of each request is text the model reads again: the whole conversation is sent with every request. Anthropic's simulation of a session, read on 4 October 2026, shows the proportions. Seven items that add up to 7,850 tokens load before the user types, the user's prompt is 45 tokens, and the four files Claude reads to answer it add 6,900 tokens. Each later request contains all of that text again.

What should I change first to lower Claude Code usage?

Change your habit between tasks first: run `/rename` and then `/clear` each time you start unrelated work in Claude Code. Anthropic's documentation, read on 4 October 2026, says `/clear` costs nothing and that old context is sent again with every later message. The second change is the model. The default on Anthropic's plans is Opus 5.5, and Anthropic advises Sonnet for most coding tasks.

Does prompt caching reduce Claude Code costs?

Yes. Prompt caching reduces Claude Code costs, and Claude Code uses it without any setup. The prompt cache is a store of request text the service has already processed, and text read from it is billed at a lower price. On Anthropic's pricing page, read on 4 October 2026, a cache read on Claude Sonnet 5.5 is listed at $0.20 per million tokens against $2 for new input. Switching models in the middle of a task starts a new cache.

Do long sessions use up a Claude subscription limit faster?

Yes. A long Claude Code session uses a subscription's limits faster than a short one doing the same work. Anthropic's documentation, read on 4 October 2026, says a session that has been open for hours can use far more of your plan limits than your activity suggests, because a one-line question still counts usage for the whole conversation. On a paid plan, `/usage` flags long context when it accounts for 10 percent or more of recent usage.

Does Claude Code charge per token on a Pro or Max subscription?

On a Pro or Max subscription, Claude Code usage is included in the plan, and it counts against the plan's usage limits. Anthropic's documentation, read on 4 October 2026, says the dollar figure in the Session block of `/usage` is intended for API users and has no bearing on a subscriber's billing. Usage credits, if you turn them on, let you keep working past the plan's limit.

How do I estimate Claude Code costs for a team?

Estimate Claude Code costs for a team by measuring a small test group first. Anthropic's documentation, read on 4 October 2026, advises starting with a small group and using its tracking tools to measure usage before giving the tool to everyone. Its published average of $13 per developer per active day is a first comparison. On a Team or Enterprise plan the spend report shows estimated spend for each user, and on the Claude Console the dashboard shows spend for each member.

Does a smaller CLAUDE.md lower Claude Code costs?

Yes. A smaller CLAUDE.md lowers Claude Code costs, because the file loads at the start of every session and is then part of every request. Anthropic's documentation, read on 4 October 2026, sets a target of under 200 lines for each file and advises moving instructions for specific procedures into skills, which load only when used. The effect is smaller than clearing the conversation, since the conversation has no fixed size.

Can an administrator limit Claude Code spending?

Yes. An administrator can limit Claude Code spending in the place that matches how the organisation pays. Anthropic's documentation, read on 4 October 2026, says organisations on the Claude Console, its website for pay-per-token accounts, set workspace spend limits, and Team and Enterprise plans set spend limits on usage credits for the organisation, a group or one member. On Amazon Bedrock, Google Cloud or Microsoft Foundry, the limits are in the cloud provider's own billing controls.

Does the 1 million token context window cost more per token?

The 1 million token context window in Claude Code is billed at the standard price per token. Anthropic's documentation, read on 4 October 2026, says it uses standard model pricing with no extra charge for tokens beyond 200,000, and the pricing page gives the example that a 900,000-token request is billed at the same rate per token as a 9,000-token request. A larger window still means larger requests if the conversation is never cleared.