Checklist

Claude Code token checklist: a one-page list to print

This Claude Code token checklist lists the actions that reduce token usage, grouped by when you do them: before a session, during a session, between tasks, once per project and once per team. Each item names the command or setting, such as /model, /context, /clear, /rewind or a CLAUDE.md file under 200 lines, and gives the reason from Anthropic's documentation, read on 4 October 2026. The first item to adopt is running /clear between unrelated tasks, which Anthropic says costs nothing. A summary table at the top lists the main items in one place, and each item links to the page of this guide that explains it.

Published October 4, 2026. Editorial.

Key takeaways

  • Running `/rename` and then `/clear` between unrelated tasks is the first item to adopt, because Anthropic's documentation, read on 4 October 2026, says `/clear` costs nothing and names uncleared long sessions as a usual cause of high spending.
  • Choose the model and the effort level before the first request, because each model has its own prompt cache and a `/model` switch makes the next request read the entire conversation with no cache reads.
  • The prompt cache lasts one hour on a Claude subscription and five minutes on usage credits, a pay-per-token API key or a cloud provider by default, so the first request after a longer break processes the full context again.
  • Anthropic's target for each CLAUDE.md file is under 200 lines, and its example hook passes at most 100 lines of failing test output to the model.
  • Anthropic states an average of $13 per developer per active day across enterprise deployments (large organisations), and advises measuring a small test group before setting a team budget.

This is the working checklist for the guide on how to reduce Claude Code token usage. Print it and keep it beside your keyboard. Each item is one action, with the command or setting that performs it, the reason in one or two sentences, and a link to the page that explains it in full.

Three words appear in almost every item, so here they are once. A token is a piece of text that the model processes, and it is the unit Claude Code usage is counted in. The context window is all the text the model reads in one request: Anthropic's instructions, your instruction files, the whole conversation so far, and every file and command result Claude has collected. The prompt cache is a store of request text that the service has already processed, and text read from it is billed at a lower price.

One fact explains the whole list. Anthropic's documentation, read on 4 October 2026, says that Claude Code sends your full conversation with every request [1]. Almost every item below either makes that conversation shorter or keeps the cache in use. The page on where Claude Code tokens go shows what a request contains.

The checklist in one table

Every command and setting in this table comes from Anthropic's documentation, read on 4 October 2026. The sections after the table give the reason for each row. Four names in the table are explained in those sections: CLAUDE.md (your instruction file), a subagent (a second copy of Claude with its own context window), a hook (a script that Claude Code runs by itself) and compaction (replacing the conversation with a summary).

Action Command or setting When
Choose the model and the effort level /model sonnet, /effort Before a session
See what loads at the start /context Before a session
Turn off outside tool connections you are not using /mcp Before a session
Name the file, the function and the check Your own request During a session
Plan before a large change Shift+Tab or /plan During a session
Stop a wrong approach, then remove it Esc, then /rewind During a session
Give long output and open-ended reading to a subagent Ask in your request During a session
Summarise at a pause in one long task /compact with instructions During a session
Keep the model for the whole task No /model change During a session
Name the session, then start a new conversation /rename, then /clear Between tasks
Cancel repeating tasks Ask Claude to list and cancel them Between tasks
Read the usage breakdown /usage Between tasks
Close helper sessions you no longer need Ask each teammate to shut down Between tasks
Keep CLAUDE.md under 200 lines Edit the file, /memory Once per project
Tell compaction what to keep A # Compact instructions section Once per project
Filter test output A PreToolUse hook Once per project
Prefer command-line tools gh, aws, gcloud, sentry-cli Once per project
Put simple subagent work on Haiku model: haiku Once per project
Measure a baseline with a small group /usage, the Console usage page Once per team
Set spending limits Admin settings or workspace limits Once per team
Limit the effort level maxEffortLevel Once per team

Before a session

  • Choose the model and the effort level first. Run /model sonnet unless the task is an architecture decision. Anthropic's documentation says Sonnet handles most coding tasks well and costs less than Opus, and the default on Anthropic's plans is Opus 5.5 [1][4]. On Amazon Bedrock and Google Cloud's Agent Platform, the name sonnet selects Sonnet 4.5 [4]. The effort level is the setting for how much the model reasons before it replies, and that reasoning is billed as output. Choose both now, because each model has its own cache and a change in the middle of a task makes the next request process the whole conversation again [2]. See which model and effort level to use.
  • Run /context once in a new session. It shows current usage by category, including which instruction files loaded [3]. Everything listed there is sent with every request you make in the session, so a large category is the first thing to reduce. See how to read /usage and /context.
  • Run /mcp and disable servers you are not using. An MCP server is a program that connects Claude Code to an outside tool or data source. By default each one adds its tool names and its instructions to the start of the session [1]. See MCP servers or command-line tools.
  • Start a new conversation for a new task. If the previous task is still in the session, its text will be sent with every request of the new one. The item on /clear below gives the steps.

During a session

  • Name the file, the function, the change and the check. Anthropic's own example of a specific request is "add input validation to the login function in auth.ts" [1]. A request that names the file removes the search, and each file Claude does not read is a file that is not sent again on every later request. See how to write a request that reads fewer files.
  • Use plan mode before a large change. Plan mode is a setting in which Claude reads and proposes an approach, and edits nothing until you approve. Press Shift+Tab to reach it, or begin the request with /plan [7]. A wrong approach found in a plan costs one correction. Found after the edits, it costs the edits too.
  • Press Esc when the approach is wrong, then run /rewind. Esc stops Claude in the middle of a turn [8]. /rewind returns the conversation to an earlier prompt, which removes the mistaken steps from every later request, and the next request reads text the cache already holds [2]. See /clear, /compact or /rewind.
  • Give long output to a subagent. A subagent is a second copy of Claude that works on one task in its own separate context window and returns a summary. Test runs, log files and documentation pages are the cases Anthropic names [1]. In Anthropic's simulation of a session, a subagent reads 6,100 tokens of files and returns 420 tokens to the main conversation [3]. The subagent's own requests still count in your usage [1].
  • Run /compact at a pause, with instructions. Compaction replaces the conversation history with a summary. Type what to keep, as in Anthropic's example /compact Focus on code samples and API usage [1]. Choose the moment yourself: Anthropic advises running it at a pause in the work, such as between tasks, instead of waiting for automatic compaction in the middle of a task [2]. Writing the summary is one request that reads the conversation, so compact a long task and clear between unrelated tasks.
  • Keep the same model until the task ends. A switch with /model means the next request reads the entire conversation with no cache reads [2]. If the task needs a different model, change after /clear.
  • Finish the step before a long break. The cache lasts one hour on a subscription. It lasts five minutes when you are using usage credits, which are paid usage beyond your plan, and by default on an API key (a pay-per-token account) or a cloud provider [1]. The first request after a longer break processes the full context again [1]. See why usage keeps rising in a long session.

Between tasks

  • Run /rename, then /clear. /clear starts a new conversation with an empty context, and Anthropic's documentation says it costs nothing [1]. Naming the session first lets you find it again with /resume [1]. This is the item to keep if you keep only one: the documentation names "long sessions that were never cleared" as a usual cause of unexpectedly high spending [1].
  • Cancel repeating tasks before you leave. A scheduled task, such as one started with /loop, runs on its interval while the session is idle and sends the full context each time [1]. Ask Claude in plain words to list your scheduled tasks and cancel the ones you do not need. A repeating task that you forget ends by itself seven days after it was created [6].
  • Check for a goal that is still active. A goal is a finish condition that keeps Claude working without a prompt from you. While background work keeps a goal waiting, Claude Code can start a new turn in an idle session to check on it, at most three times between your prompts from Claude Code v2.1.246. Setting CLAUDE_CODE_GOAL_CHECKIN_MINUTES to 0 turns these checks off [1].
  • Read the /usage breakdown. On a Pro, Max, Team or Enterprise plan, /usage flags a behaviour, such as long context or cache misses, when it accounts for 10 percent or more of your recent usage. Press d or w to change between the last 24 hours and the last 7 days [1]. The flag tells you which section of this checklist to read again.
  • Shut down agent teammates. An agent team is a group of separate Claude Code sessions that work together. Each active teammate keeps using tokens until it exits or the session ends [1]. Anthropic puts the usage of a team at 7 times a standard session when the teammates run in plan mode [1].

Once per project

  • Keep CLAUDE.md under 200 lines. CLAUDE.md is the instruction file that Claude reads at the start of every session, so every line is sent with every request. Anthropic's target is under 200 lines for each file [5]. Splitting the file into imported files, which are other files that CLAUDE.md includes by reference, leaves the cost the same, because imported files load at the start too [5]. See how long CLAUDE.md should be.
  • Move procedures into skills. A skill is a packaged set of instructions that loads only when it is used. Steps for one job, such as a code review or a database change, are the content Anthropic advises moving out of CLAUDE.md [1].
  • Add a section named # Compact instructions to CLAUDE.md. It tells every compaction in the project what to keep, such as test output and code changes [1]. You write it once, and each developer no longer has to type the instruction.
  • Install a hook that filters test output. A hook is a script that Claude Code runs by itself at a fixed point. Anthropic's example hook rewrites a test command so that only the lines containing FAIL, ERROR or error: reach the model, with five lines after each and at most 100 lines in total [1]. Check that your test tool prints one of those words on a failure, and confirm the hook is listed by running /hooks [1]. See hooks that trim test and log output.
  • Prefer command-line tools to MCP servers where both exist. Anthropic names gh, aws, gcloud and sentry-cli, and says they use less context because they add no listing of tools [1].
  • Install a code intelligence plugin for a typed language. By Anthropic's account, one "go to definition" call replaces a text search followed by reading several candidate files [1].
  • Set model: haiku on subagents that do simple work. Without a setting, a subagent runs on the main conversation's model, so a session on Opus runs its subagents on Opus [1][4].

Once per team

  • Measure before you set a budget. Anthropic states an average of $13 per developer per active day across enterprise deployments (large organisations), and $150 to $250 per developer per month, with 90 percent of users below $30 per active day [1]. Those are Anthropic's averages for its own customers. Anthropic's advice for your own estimate is to start with a small test group and record a baseline, which is a measurement taken before any change, before giving the tool to the whole team [1].
  • Set spending limits where your plan allows them. On the Claude Console, which is Anthropic's website for pay-per-token accounts, an organisation uses the control named workspace spend limits. On a Team or Enterprise plan, each member has a usage allowance, and an administrator who turns on usage credits sets spend limits for the organisation, a group or one member [1]. On a cloud provider such as Amazon Bedrock or Google Cloud's Agent Platform, usage is billed per token to your cloud account, and the spending controls are in the cloud provider's billing console [1].
  • Limit the effort level for the organisation. The maxEffortLevel setting, which an administrator applies for the whole organisation, sets the highest effort level a developer can choose, on any plan and any provider [4]. On an Enterprise plan, an administrator can also set the organisation's default model, which requires Claude Code v2.1.196 or later [4].
  • Teach two habits before any setting. Anthropic's documentation names the two habits with the highest impact: clearing between unrelated tasks, and matching the model to the job [1]. Both need no setting and no administrator.
  • Leave agent teams off unless a task needs them. They are off by default, and the setting CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1 turns them on [1]. Reveneau recommends making that a team decision, given the 7 times figure above.

What needs no action

Two things use tokens and are too small to spend time on. Background work, such as the summaries Claude Code writes so that a session can be resumed, is typically under $0.04 per session by Anthropic's figure [1]. Prompt suggestions, the next requests that Claude Code proposes after a response, reuse the cache and add a few output tokens each [1]. Turn them off if you never use them, and expect a small saving.

How to tell whether the checklist is working

Compare two normal working days: one before you adopt the list, and one after. Three numbers from /usage are enough.

  • The token counts by model. The Session block lists input, output, cache read and cache write for each model [1]. In a session where the cache is in use, most of the input appears as cache reads, because most of each request repeats the request before it [2].
  • The prompt cache line. From Claude Code v2.1.251, the Session block has a line named Prompt cache (main) with the share of input tokens read from the cache and the number of misses [1]. The related guide on prompt caching in Claude Code explains how to read it.
  • The behaviour flags. If long context or cache misses still appear as flags after a week, the habits under "Between tasks" are the ones that are being skipped.

These figures are estimates that Claude Code computes on your computer. For billing, Anthropic's documentation names the Usage page in the Claude Console as the authoritative record [1].

About this checklist

Reveneau is an AI software development consultancy, and all of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau recommends this order: the habits between tasks first, the project setup second, and the team settings third. Reveneau is independent of Anthropic. Every command, setting, version number and figure on this page comes from Anthropic's own documentation, read on 4 October 2026. Commands change between versions, so run claude --version and check the documentation when a command on this list does not behave as described [1].

Common questions

What should I check before starting a Claude Code session?

Before starting a Claude Code session, check three things: the model, what loads at the start, and which outside tool servers are connected. Run `/model sonnet` unless the task is an architecture decision, run `/context` to see usage by category, and run `/mcp` to disable servers you are not using. Anthropic's documentation, read on 4 October 2026, says each model has its own prompt cache, so the model is best chosen before the first request.

In what order should I adopt the Claude Code token checklist?

Adopt the Claude Code token checklist in three stages: the habits between tasks first, the project setup second, and the team settings third. That is the order Reveneau recommends. The first stage starts with running `/rename` and then `/clear` between unrelated tasks. Anthropic's documentation, read on 4 October 2026, says `/clear` costs nothing and names long sessions that were never cleared as a usual cause of unexpectedly high spending.

How often should I run /clear in Claude Code?

Run `/clear` in Claude Code each time you change to a task that has no use for the current conversation. Anthropic's documentation, read on 4 October 2026, advises it when switching to unrelated work, because old context is sent again with every later message. Run `/rename` first, so that you can return to the old conversation with `/resume`. Inside one long task, use `/compact` at a pause, since that keeps a summary.

What should I set up once per project to lower Claude Code usage?

Set up five things once per project to lower Claude Code usage: a CLAUDE.md file under 200 lines, skills for step-by-step procedures, a `# Compact instructions` section, a hook that filters test output, and `model: haiku` on subagents that do simple work. Each comes from Anthropic's documentation, read on 4 October 2026. They take effort once and then apply to every session of every developer on the project.

What should a team set up once to control Claude Code costs?

A team should set up a first measurement, spending limits and an effort limit to control Claude Code costs. Anthropic's documentation, read on 4 October 2026, advises starting with a small test group and measuring its usage before giving the tool to the whole team. Organisations on the Claude Console set workspace spend limits, and an administrator's `maxEffortLevel` setting limits the effort level on any plan. Agent teams are off by default, and Reveneau recommends making it a team decision to turn them on.

Should my team turn on agent teams in Claude Code?

Turn on agent teams in Claude Code only when a task needs them, and make that a team decision. That is Reveneau's recommendation. Anthropic's documentation, read on 4 October 2026, says agent teams are off by default and that the setting `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1` turns them on. Anthropic puts the usage of a team at 7 times a standard session when the teammates run in plan mode, and each active teammate keeps using tokens until it exits.

What should I do during a Claude Code session to keep token usage low?

During a Claude Code session, keep token usage low with five habits: name the file and the check in each request, use plan mode before a large change, press `Esc` and run `/rewind` when the approach is wrong, give long output to a subagent, and keep the same model until the task ends. Anthropic's documentation, read on 4 October 2026, says a switch with `/model` makes the next request read the entire conversation with no cache reads.

How do I print the Claude Code token checklist?

Print the Claude Code token checklist with the button labelled "Print or save as PDF" on this page, or with your browser's print command. The summary table near the top lists each action with its command and the moment to use it, and it is the short version to print. The checklist reflects Anthropic's documentation as read on 4 October 2026, so check the date before you rely on an old printed copy.

Which Claude Code commands matter most for token control?

Six Claude Code commands do most of the work of token control: `/clear`, `/compact`, `/rewind`, `/context`, `/usage` and `/model`. By Anthropic's documentation, read on 4 October 2026, `/clear` starts a new conversation with an empty context, `/compact` replaces the history with a summary, and `/rewind` returns to an earlier prompt. `/context` and `/usage` measure, and `/model` sets the price of each token.

How do I know whether the token checklist is working?

You know the token checklist is working when the `/usage` figures for a normal day go down and the behaviour flags stop appearing. Compare the token counts by model before and after. From Claude Code v2.1.251, by Anthropic's documentation read on 4 October 2026, the Session block also shows a `Prompt cache (main)` line with the share of input tokens read from the cache. The dollar figure there is an estimate computed on your computer.

Does the Claude Code token checklist apply on Amazon Bedrock or Google Cloud?

Yes, the Claude Code token checklist applies on Amazon Bedrock and Google Cloud's Agent Platform, with two differences. Anthropic's documentation, read on 4 October 2026, says the prompt cache lasts five minutes by default on a cloud provider, so a break of a few minutes is enough for the cache to expire. It also says the `sonnet` alias selects Sonnet 4.5 on those two providers. Usage there is billed per token to your cloud account, and spending limits are set in the cloud provider's billing console.

What can I ignore when reducing Claude Code token usage?

You can ignore background usage when reducing Claude Code token usage. Anthropic's documentation, read on 4 October 2026, says background work such as conversation summaries typically costs under $0.04 per session. Prompt suggestions are mostly cache reads plus a few output tokens. The large amounts are in the conversation itself: its length, the files and command output in it, and the model that reads it on every request.