Start here

Where Claude Code tokens go: what is sent on every message

Claude Code tokens go mostly to text the model reads again on every request. Each message re-sends the whole conversation: Anthropic's instructions, your CLAUDE.md instruction files, every earlier message, and every file and command result Claude has collected. In Anthropic's own simulation, read on 4 October 2026, the seven items that load before the user types add up to 7,850 tokens, and four file reads then add 6,900 tokens to a 45-token prompt. Input, output and cached tokens are billed at different prices, and the model's reasoning is billed as output. This page lists what is in the context window, when each part loads, and how to make each part smaller.

Published October 4, 2026. Editorial.

Key takeaways

  • Claude Code sends the full conversation with every request, because the model keeps nothing between requests, by Anthropic's documentation read on 4 October 2026.
  • In Anthropic's simulation of a session, seven items load before the user types and total 7,850 tokens, of which the system prompt is 4,200.
  • In the same simulation, a 45-token prompt leads to four file reads of 6,900 tokens in total, and the terminal shows one line for each read.
  • Thinking tokens are billed as output tokens, and Claude Sonnet 4.6 output is listed at $15 per million tokens against $3 for input.
  • A cache read on Claude Sonnet 4.6 is listed at $0.30 per million tokens, one tenth of the input price, on Anthropic's pricing page read on 4 October 2026.

Claude Code is Anthropic's coding tool. You type a request in a terminal (the text window where you type commands), and an AI model reads files, runs commands and edits code for you. All of that work is counted in tokens. Anthropic's pricing page, read on 4 October 2026, defines tokens as "pieces of text that models process" and gives an estimate of 4 characters, or 0.75 of an English word, for one token [6].

This page shows where those tokens go. It is the first page of the guide on how to reduce Claude Code token usage, because every saving in that guide follows from one fact: most of the tokens you pay for are text the model reads again on every request, and you typed almost none of it.

Every message sends the whole conversation again

The model keeps nothing between two requests. Anthropic's documentation, read on 4 October 2026, says it in one sentence: "The model doesn't remember anything between requests, so Claude Code re-sends the full context: the system prompt, your project context, every prior message and tool result, and your new message." [3]

The place that holds all of this text is called the context window. It is the largest amount of text the model can read in one request. Anthropic's interactive simulation of a session uses a window of 200,000 tokens, and the same page says that some models support a window of 1 million tokens [2].

One message from you is often several requests. A tool is an action the model can ask Claude Code to perform, such as reading a file or running a command. The documentation says that "each time Claude uses tools it sends another request carrying that batch of tool results" [1]. A request that makes Claude read five files in five steps can therefore send the conversation five times, and each time it is longer than before.

The same documentation gives the result: a one-line question, asked in a session that has been open all day, is still counted as usage for the whole conversation [1]. The size of a request is set by everything that came before your question.

What loads before you type anything

A new session already holds text before your first word. Anthropic's simulation lists seven items that load at the start, each with a token count. Anthropic calls these "representative token counts", so they are examples from Anthropic's own illustration, and your session will differ [2].

What loads at the start Size in Anthropic's simulation What it holds
System prompt 4,200 tokens Anthropic's core instructions for behaviour, tool use and response format
Project CLAUDE.md 1,800 tokens Your project's conventions, build commands and notes on how the code is organised
Auto memory (MEMORY.md) 680 tokens Notes Claude wrote for itself in earlier sessions
Skill descriptions 450 tokens One line for each available skill
User CLAUDE.md in your home folder 320 tokens Your personal preferences for every project
Environment information 280 tokens The folder you are working in, the type of computer and its operating system version
MCP tool names 120 tokens The names of tools from connected outside services

Added together, the seven items come to 7,850 tokens. That sum is our own arithmetic on Anthropic's figures, and it equals 3.9 percent of the 200,000-token window in the simulation. Each term in the table needs one plain explanation.

The system prompt is the set of instructions Anthropic writes for the model. You never see it, and it always loads first [2]. It also contains the tool definitions, which are the descriptions that tell the model what each tool does and how to call it [3].

CLAUDE.md is a text file of instructions that you write and that Claude reads at the start of every session. Claude Code loads the file from the folder you start it in and from every folder above that one, and joins them together [4]. The documentation states the cost directly: files over 200 lines "consume more context and may reduce adherence" [4]. The page on how long CLAUDE.md should be covers what to keep in it.

Auto memory is a set of notes that Claude writes for itself about your preferences and corrections. The first 200 lines of its index file, MEMORY.md, or the first 25KB, whichever comes first, load at the start of every conversation [4].

A skill is a packaged set of instructions for one task. Only a one-line description of each skill loads at the start. In the words of the simulation, "Full skill content loads only when Claude actually uses one." [2]

An MCP server is a program that connects Claude Code to an outside tool or data source, such as an issue tracker. MCP stands for Model Context Protocol. By default Claude Code uses a feature named tool search, which holds back the full definition of each MCP tool until Claude needs it. Anthropic's documentation says that "only tool names and server instructions load at session start" [5]. If tool search is turned off with ENABLE_TOOL_SEARCH=false, every definition loads at the start [5]. The page on MCP servers and command-line tools compares the two.

All of this start-up text is sent again with every request for the whole session. A line you add to a project's CLAUDE.md is sent on every request, in every session, for every person who uses that project.

What each file read and tool result adds

After the start, the largest additions come from Claude's own work. The simulation follows one bug fix. Your prompt in it is 45 tokens. Claude then reads four files of 2,400, 1,100, 1,800 and 1,600 tokens, which is 6,900 tokens in total, or 153 times the size of the prompt (our arithmetic on Anthropic's illustrative figures) [2]. Anthropic's own note beside the first read says: "File reads dominate context usage." [2]

Your terminal hides most of this. The simulation explains that you see a single line such as "Read auth.ts", while the 2,400 tokens of file content go only to the model [2]. The same holds for command output. A search across the code adds 600 tokens in the simulation, and one run of the project's automated tests adds 1,200 tokens, while your screen shows a short status line [2].

Four smaller additions arrive along the way, by Anthropic's account on 4 October 2026:

  • Path rules. A rule file can name the file paths it applies to. It loads when Claude first reads a matching file. The simulation shows two such rules, of 380 and 290 tokens [2].
  • Hook output. A hook is a script that Claude Code runs by itself at a fixed point, for example after every file edit. Text that the hook sends back for Claude to read enters the context. The simulation shows 120 tokens from one hook that formats code after an edit [2].
  • Commands you run yourself. When you type ! before a command, Claude Code runs that command for you. The command and its output both enter the context as part of your message [2].
  • Notices about changed files. If a file changes after Claude read it, the earlier copy stays in the history. Claude Code adds a notice, and Claude reads the file again if needed [3]. The file can then be in the conversation twice.

A subagent changes where this text goes. A subagent is a second copy of Claude that works on one task in its own separate context window and sends back a summary. In the simulation, a subagent reads 6,100 tokens of files and returns a result of 420 tokens to the main conversation [2]. The main conversation stays small. The subagent's own requests are still counted in your usage, as the documentation states [1].

Input, output and cached tokens have different prices

Claude Code reports four counts for each model: input, output, cache read and cache write [1]. Each has its own price.

Input tokens are text the model reads. Output tokens are text the model writes, which includes its replies and its reasoning.

Cache read and cache write tokens come from the prompt cache. The prompt cache is a store, kept by the service that runs the model, of request text it has already processed. When the start of a new request matches the stored text exactly, the service reads that part from the store at a lower price and fully processes only the new part [3]. A cache write is the first storing of a piece of text. A cache read is each later reuse.

The table gives Anthropic's list prices for one model, Claude Sonnet 4.6, from Anthropic's pricing page read on 4 October 2026 [6]. Prices for other models are on the same page.

Token type What it is List price for Claude Sonnet 4.6, per million tokens
Input New text the model reads at full price $3
Cache write, five-minute store Text stored for reuse for the first time $3.75
Cache write, one-hour store The same, kept for longer $6
Cache read Stored text read again $0.30
Output Text the model writes, including reasoning $15

Two ratios in that table matter. A cache read costs one tenth of the input price, so the stored part of a long conversation costs one tenth of what the same text would cost if it were processed in full. Output costs five times the input price on this model, so long replies and long reasoning are the most expensive tokens.

Reasoning adds to the output count. Before it answers, the model can write out its reasoning, which Anthropic calls extended thinking. The documentation, read on 4 October 2026, says that extended thinking is on by default, that "Thinking tokens are billed as output tokens", and that the default allowance "can be tens of thousands of tokens per request depending on the model" [1]. One control is the effort level, a setting for how much the model reasons before it replies. The same documentation says thinking cannot be turned off on Opus 5.5, Sonnet 5.5 or the Fable models [1]. The page on which model and effort level to use covers that choice.

The cache has one strict rule: it matches from the start of the request, and a change early in the request means everything after it is processed again in full [3]. The related guide on prompt caching explains the mechanism, starting with what the prompt cache is.

What is in the context window, when it loads, and how to make it smaller

This table summarises the page. Each fix in the last column comes from Anthropic's documentation, read on 4 October 2026.

Content When it loads How to make it smaller
System prompt and tool definitions First, on every request Outside your control, apart from the set of tools: run /mcp and disable servers you are not using [1]
Project and user CLAUDE.md At the start of the session Keep it under 200 lines and move step-by-step procedures into skills [1]
Auto memory At the start: the first 200 lines or 25KB Run /memory to read, edit or delete the notes [4]
Skill descriptions At the start, one line per skill A skill marked disable-model-invocation: true stays out of the list until you call it [2]
MCP tool names At the start Prefer command-line tools such as gh, which add no tool listing [1]
Path rules and nested CLAUDE.md files When Claude reads a matching file These already load only when needed [4]
File contents Each time Claude reads a file Name the file and the function in your request [1]
Command and tool output Each time a tool runs Filter it with a hook, or give the task to a subagent [1]
Conversation history Grows with every turn Run /clear between unrelated tasks [1]
Replies and reasoning Each response, billed as output Lower the effort level for simple tasks [1]

Where to start

Our position: shorten the conversation history first, and the start-up text second. The start-up text in Anthropic's simulation is 7,850 tokens and stays that size. The history has no upper limit except the window, and all of it is sent again on every request. Reveneau recommends one habit before any edit to a settings file: run /clear when you change to an unrelated task. Anthropic's documentation names "long sessions that were never cleared" as a usual cause of unexpectedly high spending [1].

Measure before you change anything. The page on how to read /usage and /context shows which screen gives which number, and why usage keeps rising in a long session lists the eight causes Anthropic documents.

Reveneau is an AI software development consultancy, and all of its code is written by AI, so token use is a running cost of every Reveneau build. Reveneau is independent of Anthropic. Every figure on this page is Anthropic's own statement about its own product.

Common questions

What is a token in Claude Code?

A token in Claude Code is a piece of text that the model processes, and it is the unit all usage is counted in. Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters or 0.75 of an English word, and says the exact count varies by language and content. File contents, command output and the model's replies are all counted in tokens.

Does Claude Code send the whole conversation with every message?

Yes. Claude Code sends the whole conversation with every request, because the model keeps nothing between requests. Anthropic's documentation, read on 4 October 2026, lists what is sent again: the system prompt, your project context, every earlier message and tool result, and your new message. When Claude uses tools, each batch of tool results goes in a further request that contains the conversation again.

How many tokens does Claude Code use before I type anything?

The number of tokens Claude Code loads before you type depends on your setup, and Anthropic's own simulation shows 7,850. That simulation, read on 4 October 2026, lists seven start-up items: a 4,200-token system prompt, 1,800 tokens of project CLAUDE.md, and smaller amounts for memory notes, skill descriptions and MCP tool names. Anthropic calls these representative counts. Run `/context` to see your real figures.

Why does reading a file use so many tokens in Claude Code?

Reading a file uses many tokens because the full content of the file enters the conversation and is then sent again with every later request. In Anthropic's simulation, read on 4 October 2026, one file read adds 2,400 tokens while the terminal shows a single line. Four file reads in that simulation add 6,900 tokens to a prompt of 45 tokens. Naming the exact file in your request reduces the number of reads.

Are thinking tokens billed in Claude Code?

Yes. Thinking tokens in Claude Code are billed as output tokens, the token type with the highest list price. Anthropic's documentation, read on 4 October 2026, says extended thinking is on by default and that the default allowance can be tens of thousands of tokens per request, depending on the model. Lowering the effort level, the setting for how much the model reasons, reduces this cost on simple tasks.

What is the difference between input tokens and output tokens?

Input tokens are text the model reads, and output tokens are text the model writes. Output has the higher price. On Anthropic's pricing page, read on 4 October 2026, Claude Sonnet 4.6 is listed at $3 per million input tokens and $15 per million output tokens, a ratio of five to one. The model's reasoning, which Anthropic calls extended thinking, is billed as output tokens.

What are cache read and cache write tokens?

Cache read and cache write tokens are input tokens that pass through the prompt cache, a store of request text the service has already processed. A cache write is the first storing of a piece of text, and a cache read is each later reuse. On Anthropic's pricing page, read on 4 October 2026, a cache read costs 0.1 times the input price on most models, and a five-minute cache write costs 1.25 times.

Does command output use tokens when my terminal shows only one line?

Yes. Command output uses tokens in full in Claude Code, even when the terminal shows a short status line. In Anthropic's simulation, read on 4 October 2026, one search across the code adds 600 tokens and one run of the tests adds 1,200 tokens, while the screen shows a short line. File reads work the same way: the terminal shows one line such as "Read auth.ts", and the 2,400 tokens of file content go only to the model.

Do commands I run with the ! prefix add tokens in Claude Code?

Yes. A command you run with the `!` prefix adds tokens in Claude Code, because the command and its output both enter the context as part of your message. That is the description in Anthropic's simulation, read on 4 October 2026. The output then stays in the conversation and is sent again with every later request, in the same way as a file Claude has read. Running `/clear` when you change to unrelated work removes it.

Does a subagent reduce token usage in Claude Code?

A subagent reduces the tokens in your main conversation, and its own work is still counted in your usage. A subagent is a second copy of Claude with a separate context window. In Anthropic's simulation, read on 4 October 2026, a subagent reads 6,100 tokens of files and returns a 420-token result, so the main conversation grows by 420 tokens. Later requests in the main conversation then include 420 tokens in place of 6,100.

Can I see the text Claude Code sends to the model?

You can see a breakdown of that text by category with the `/context` command, while much of the text itself stays hidden in the terminal. Anthropic's simulation, read on 4 October 2026, marks the system prompt, memory notes and skill descriptions as invisible in your terminal, and shows each file read as one line. `/context` reports current usage by category, including which CLAUDE.md and auto memory files loaded.

Which part of the context window should I reduce first?

Reduce the conversation history first, because it is the part of the context window that keeps growing. The start-up content in Anthropic's simulation is 7,850 tokens and stays that size, while each file read and command result adds to the history and is sent again on every request. Anthropic's documentation, read on 4 October 2026, advises running `/clear` when you change to unrelated work.