How to reduce Claude Code token usage / Choices
How to write a Claude Code request that reads fewer files
To write a Claude Code request that reads fewer files, name the file, the function, the change you want, and the check that shows the work is done. Anthropic's documentation, read on 4 October 2026, says a vague request such as "improve this codebase" causes broad scanning, while a request that names a function and a file needs few file reads. In Anthropic's simulation of a session, a 45-token prompt that names no file leads to four file reads of 6,900 tokens in total. This page gives examples of vague and specific requests, and covers plan mode, stopping early with Escape and /rewind, test targets, and giving exploration to a subagent.
Published October 4, 2026. Editorial.
Key takeaways
- Anthropic's documentation, read on 4 October 2026, contrasts the vague request "improve this codebase" with the specific request "add input validation to the login function in auth.ts".
- In Anthropic's simulation of a Claude Code session, a 45-token prompt that names a symptom and no file leads to four file reads of 2,400, 1,100, 1,800 and 1,600 tokens.
- Plan mode in Claude Code starts with `Shift+Tab`, the `/plan` prefix or `claude --permission-mode plan`, and Claude edits no source file until you approve the plan.
- Pressing `Esc` stops Claude in the middle of a turn, and `/rewind` then removes the mistaken steps from the conversation, with file copies kept for the 100 most recent checkpoints.
- In Anthropic's simulation a subagent reads 6,100 tokens of files and returns a 420-token result, and the subagent's own requests still count in your usage.
Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. All of that work is counted in tokens. A token is a piece of text that the model processes, and it is the unit your usage is measured in.
The request you type can be a small part of that usage. In Anthropic's own simulation of a session, read on 4 October 2026, the user's prompt is 45 tokens, and the four files Claude reads to answer it add 6,900 tokens [2]. The wording of a request decides how much searching Claude has to do. This page, part of the guide on how to reduce Claude Code token usage, shows how to write a request that leads to fewer reads, and what to do when Claude starts on the wrong approach.
Why a vague request costs more than a specific one
Anthropic's documentation, read on 4 October 2026, gives the rule with its own example: "Vague requests like "improve this codebase" trigger broad scanning. Specific requests like "add input validation to the login function in auth.ts" let Claude work efficiently with minimal file reads." [1]
The reason is what happens to a file after Claude reads it. The full content of the file enters the context window, which is all the text the model reads in one request. It stays there. Claude Code sends the whole conversation again with every request, and "each time Claude uses tools it sends another request carrying that batch of tool results" [1]. A file that was read at the start of a task is sent again on every later step of that task. The page on where Claude Code tokens go explains this in full.
The simulation shows the pattern. Its prompt is "Fix the auth bug where users get 401 after token refresh". The prompt names a symptom and no file. Claude reads the main auth file (2,400 tokens), follows a reference to a second file (1,100 tokens), reads a third file to trace the flow further (1,800 tokens), and reads the test file to learn the expected behaviour (1,600 tokens) [2]. Anthropic's note beside the first read says: "File reads dominate context usage. Be specific in prompts ("fix the bug in auth.ts") so Claude reads fewer files." [2]
Your terminal hides the size of this. The simulation states that you see one line such as "Read auth.ts", while the 2,400 tokens of content go only to the model [2].
What a specific request names
A specific request names four things. Each one removes a search that Claude would otherwise have to do.
- The file. Typing
@in the prompt starts file path autocomplete, so you can pick the exact path without typing it in full [5]. - The function or the place in the file. Anthropic's example names "the login function".
- The change. "Add input validation" tells Claude what to write. "Improve" leaves Claude to work out what you want by reading more code.
- The check. Say how Claude can tell that the work is done, such as a test that must pass. A later section covers this.
The table compares vague and specific versions. The first row is Anthropic's example [1]. The other rows are our own invented examples, with invented file names, written to show the same rule. Two terms in the table are explained later on this page. A subagent is a second copy of Claude that works on one task in its own separate context window and sends back a summary. The /plan prefix starts plan mode, in which Claude proposes an approach before it edits anything.
| Vague request | Specific version | Why it reads less |
|---|---|---|
| "improve this codebase" | "add input validation to the login function in auth.ts" | It names one file and one function, so Claude has no reason to scan folders |
| "fix the login bug" | "users get a 401 error after a token refresh. The refresh code is in src/lib/tokens.ts. Run the auth tests when you finish" | It names the file and the check, so Claude has less searching to do to find the fault |
| "why are the tests failing?" | "use a subagent to run the test suite and report only the failing tests with their error messages" | The full test output stays in the subagent's separate context, and a short report comes back |
| "add caching to the API" | "/plan add caching to the getProducts function in api/products.ts" | Claude proposes an approach first, and you correct it before any file is edited |
| "clean up the utils" | "in src/utils/dates.ts, remove the two functions that no other file uses" | It limits the work to one file and one type of change |
The subagent wording in the third row is taken from an example prompt in Anthropic's documentation [6].
Use plan mode before large work
Plan mode is a setting in which Claude researches and proposes changes without making them. Anthropic's documentation, read on 4 October 2026, says that in plan mode Claude "reads files, runs shell commands to explore, and writes a plan, but does not edit your source" [3]. A shell command is a command typed into the terminal.
You can start it in three ways [3]:
- Press
Shift+Tabuntil the status bar showsplan mode on. - Begin a request with
/plan. Claude Code enters plan mode and starts on that request. - Start the session with
claude --permission-mode plan.
When the plan is ready, Claude shows it and asks how to continue. You can approve it, or choose No, keep planning and say what to change. Pressing Ctrl+G opens the plan in your text editor so that you can edit it yourself before Claude continues [3].
Plan mode reads files too, so it uses tokens. Its value is in what it prevents. Anthropic's documentation says that in plan mode Claude explores the code and proposes an approach for your approval, and that this prevents costly repeat work when the first approach is wrong [1]. A wrong approach that you find in a plan costs one short correction. The same wrong approach found after Claude has edited several files costs the edits, the reading that came before them, and the work to undo them.
Two more facts from the documentation make plan mode a low-cost habit:
- The reading can happen outside your conversation. In plan mode Claude gives research to a built-in helper named the Plan subagent, "so that exploration output stays in a separate context window" [6].
- Switching to plan mode keeps the prompt cache. The prompt cache is a store of request text that the service has already processed, and text read from it is billed at a lower price. Plan mode adds its instructions as new messages at the end of the conversation, so the stored text still matches [7]. The related guide lists the other actions that keep the cache.
One exception: with the opusplan model setting, entering or leaving plan mode changes the model, and the cache starts again [7]. The page on which model and effort level to use covers that setting.
Our position: use plan mode when you cannot name the files a task will change, or when the task will change several of them. Skip it for a change you can describe in one sentence that names the file.
Stop early with Escape, then rewind
Watch the first steps of a task. If Claude opens files that have nothing to do with your request, stop it. Pressing Esc stops the current response or tool call in the middle of the turn so that you can redirect, and Claude keeps the work done so far [5]. Anthropic's advice, read on 4 October 2026, is to press Escape at once when Claude starts on the wrong approach [1].
Stopping leaves the wrong steps in the conversation. Every file Claude read by mistake is still in the context window, and it will be sent again with each later request. The simulation makes the same point about messages: "Follow-ups add to the same context." [2] If you type "no, look in the other folder", the mistaken reads stay and your correction is added after them.
Rewinding removes them. Run /rewind, or press Esc twice when the input box is empty, to open the rewind menu [4]. The menu lists each prompt you sent in the session. You select one and choose an action [4]:
| Action in the rewind menu | What it does |
|---|---|
| Restore code and conversation | Returns both the files and the conversation to that point |
| Restore conversation | Returns the conversation to that point and keeps the current files |
| Restore code | Returns the files to that point and keeps the conversation |
| Summarize from here | Replaces the conversation from that point forward with a summary |
| Summarize up to here | Replaces the conversation before that point with a summary and keeps the later messages |
After you restore the conversation, your original prompt returns to the input box, so you can edit it into a specific request and send it again [4]. The next request then reads text that the prompt cache already holds, because the remaining conversation is the same text the cache was built from [7].
Three limits apply, by the documentation read on 4 October 2026 [4]:
- Claude Code keeps file copies for the 100 most recent checkpoints in a session. A checkpoint is the saved state before each prompt you send.
- Files changed by a shell command, such as
rmormv, are outside checkpointing. Rewind restores only edits made with Claude's own file editing tools. - Edits made by a subagent are usually not restored. Use git, the version control tool, to undo those.
The page on /clear, /compact and /rewind compares rewinding with the other two commands.
Give Claude a way to check its own work
Anthropic's documentation, read on 4 October 2026, advises: "Include test cases, paste screenshots, or define expected output in your prompt. When Claude can verify its own work, it catches issues before you need to request fixes." [1]
Here is why a checkable target reduces usage. Without one, Claude stops when the code looks finished. You then test it, find a fault, and send a new message. That message is a new turn, and the whole conversation is sent again with it. Each round of "this is wrong, fix it" repeats that cost, and the conversation is longer each time. With a target such as "the tests in auth.test.ts must pass", Claude runs the check inside the same turn and repairs a failure before it reports back. This paragraph is our reasoning from the two facts Anthropic documents: the full conversation is sent with every request, and a target lets Claude find faults before you ask [1].
A target can take three forms, all named in Anthropic's sentence: a test case, a screenshot of the expected screen, or the expected output written in the prompt [1].
This is the same idea that eval-driven development applies to a whole project. Reveneau is an AI software development consultancy, all of its code is written by AI, and every change must pass an eval suite, which is a set of automated tests written from the specification, before release. A test named in a request does for one task what the eval suite does for the build.
The same list in Anthropic's documentation adds one more habit: "Write one file, test it, then continue." [1] A fault found after one file is repaired while the conversation is still short. A fault found after many files is repaired in a conversation that already contains all of them.
Give exploration to a subagent
Sometimes you cannot name the file, because you do not know the code yet. Finding it means reading many files, and most of what is read will not matter afterwards. This is the work to give to a subagent. A subagent is a second copy of Claude that works on one task in its own separate context window and sends back a summary.
In Anthropic's simulation, a subagent reads 6,100 tokens of files and returns a result of 420 tokens to the main conversation [2]. The main conversation grows by 420 tokens. Claude Code has a built-in subagent for this, named Explore, which the documentation describes as a read-only agent for searching and analysing code [6]. You can ask for it in plain words:
# Our example request, with an invented topic
Use a subagent to find where session timeouts are handled, and report the file names and function names only.
Three cautions from the documentation, read on 4 October 2026:
- The subagent's own requests still count in your usage [1]. Delegating reduces what later requests in the main conversation contain. The subagent's own reading still uses tokens once.
- A subagent starts without your conversation history, so it may need time to collect what it needs. For "a quick, targeted change", Anthropic advises staying in the main conversation [6].
- For a question about something Claude has already read, use
/btwin place of a subagent. It answers from the conversation, reads no files, and stays out of the history [5].
Two other tools reduce exploration, and each has its own page. A skill, which is a packaged set of instructions that loads only when used, can describe how the project is organised, so that Claude learns it from the skill instead of reading many files [1]. The page on how long CLAUDE.md should be covers skills and CLAUDE.md, the instruction file that Claude reads at the start of every session. A code intelligence plugin lets Claude go directly to the place where a function is defined. The page on hooks that trim test and log output covers those plugins and hooks, which are scripts that Claude Code runs by itself at fixed points.
A request to copy
This is our own template, with invented file names. It names the file, the place, the change, the limit and the check.
# Our template, with invented file names
In src/api/auth.ts, the refreshToken function returns a 401 error after a token refresh.
Fix the order of the steps in that function. Change no other file.
Done when the tests in src/api/auth.test.ts pass.
Our position
Write the file name and the check into every request that changes code. Those two items do the most, because the first removes the search and the second removes the repeat attempts. Use plan mode when you cannot name the files. Press Esc as soon as the approach is wrong, and rewind to the prompt in place of typing a correction after it. Give open-ended reading to a subagent.
Token use is a running cost of every Reveneau build, because Reveneau uses AI instead of hiring more engineers. Reveneau is independent of Anthropic. The figures on this page come from Anthropic's own documentation and its own illustration of a session, read on 4 October 2026.
Common questions
How do I write a prompt that makes Claude Code read fewer files?
Write a prompt that names the file, the function, the change and the check, and Claude Code has less to search for. Anthropic's documentation, read on 4 October 2026, gives the example "add input validation to the login function in auth.ts" and says a request of that form needs few file reads. Add how Claude can confirm the result, such as a test that must pass, so that faults are repaired in the same turn.
Why does a vague request use more tokens in Claude Code?
A vague request uses more tokens in Claude Code because Claude has to read files to find out what you mean, and every file it reads is sent again with each later request. Anthropic's documentation, read on 4 October 2026, says a request such as "improve this codebase" causes broad scanning. In Anthropic's simulation, a prompt of 45 tokens that names no file leads to 6,900 tokens of file reads.
What is plan mode in Claude Code?
Plan mode in Claude Code is a setting in which Claude researches a task and proposes changes without making them. Anthropic's documentation, read on 4 October 2026, says that in plan mode Claude reads files, runs terminal commands to explore and writes a plan, and does not edit your source. When the plan is ready you approve it, or choose "No, keep planning" and tell Claude what to change. Approving the plan ends plan mode.
How do I start plan mode in Claude Code?
Start plan mode in Claude Code by pressing `Shift+Tab` until the status bar shows `plan mode on`, or by beginning a request with `/plan`. To open a whole session in plan mode, run `claude --permission-mode plan`. These are the three ways in Anthropic's documentation, read on 4 October 2026. Pressing `Shift+Tab` again leaves plan mode without approving a plan, and `Ctrl+G` opens a proposed plan in your text editor.
Does plan mode save tokens in Claude Code?
Plan mode saves tokens in Claude Code when it shows you a wrong approach before any file is edited. The planning itself uses tokens, because Claude still reads files. Anthropic's documentation, read on 4 October 2026, says plan mode prevents costly repeat work when the first approach is wrong. It also says Claude gives planning research to the Plan subagent, so that output stays in a separate context window.
What does pressing Escape do in Claude Code?
Pressing Escape once in Claude Code stops the current response or tool call in the middle of the turn, and Claude keeps the work done so far. Anthropic's documentation, read on 4 October 2026, advises pressing it at once when Claude starts on the wrong approach. Pressing Escape twice with an empty input box opens the rewind menu. If the input box contains text, pressing twice clears the text and saves it to your input history.
Should I correct Claude with a new message or rewind?
Rewind when the last steps were a mistake, because a new message leaves the mistaken steps in the Claude Code conversation. Every file Claude read by mistake stays in the context and is sent again with each later request. Running `/rewind` and choosing "Restore conversation" removes those steps and puts your original prompt back in the input box, by Anthropic's documentation read on 4 October 2026. You can then edit the prompt and send it again.
What is a verification target in a Claude Code prompt?
A verification target in a Claude Code prompt is something Claude can check its own work against. Anthropic's documentation, read on 4 October 2026, names three forms: a test case, a pasted screenshot, and expected output written in the prompt. It says that when Claude can verify its own work, it finds faults before you need to ask for fixes. Each fix you do not have to request is one less turn that sends the whole conversation again.
Why does testing one file at a time reduce token usage?
Testing one file at a time reduces token usage in Claude Code because a fault is found while the conversation is still short. Anthropic's documentation, read on 4 October 2026, advises: write one file, test it, then continue. A repair made early is a request that contains one file of work. A repair made after many files is a request that contains all of them, since the whole conversation is sent each time.
When should I give exploration to a subagent in Claude Code?
Give exploration to a subagent in Claude Code when finding the answer means reading many files that will not matter afterwards. In Anthropic's simulation, read on 4 October 2026, a subagent reads 6,100 tokens of files and returns a 420-token result to the main conversation. Anthropic advises staying in the main conversation for a quick, targeted change, and the subagent's own requests still count in your usage.
What does /btw do in Claude Code?
`/btw` asks a side question in Claude Code without adding it to the conversation history. Anthropic's documentation, read on 4 October 2026, says Claude answers from what is already in the conversation and has no tool access, so it reads no files and runs no commands. While the conversation's prompt cache has not expired, the question costs little beyond the answer itself. Use it to ask about code Claude has already read.
How do I tell Claude Code which file to work on?
Tell Claude Code which file to work on by writing the file's path in the request. Typing `@` in the prompt starts file path autocomplete, by Anthropic's documentation read on 4 October 2026, so you can choose the exact path from a list. Then name the function or the place inside the file. Anthropic's example request names both: the login function, in auth.ts.
References
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Explore the context window (code.claude.com), read 4 October 2026
- Anthropic, Choose a permission mode (code.claude.com), read 4 October 2026
- Anthropic, Checkpointing (code.claude.com), read 4 October 2026
- Anthropic, Interactive mode (code.claude.com), read 4 October 2026
- Anthropic, Create custom subagents (code.claude.com), read 4 October 2026
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026