Why we are publishing a Claude Code token benchmark with every raw run public

Anthropic's documentation on Claude Code, read on 4 October 2026, names the changes that reduce token usage: clear the conversation between unrelated tasks, keep CLAUDE.md short, filter long command output, choose the model and the effort level before the first message. It gives no figure for how much any of them saves. We have written five guides on those changes, and none of them carries a figure of our own, because we have measured none. That is the gap this benchmark is for. This post gives the method, fixed before the first run, and says plainly what the result will be unable to show.
Four terms first. Claude Code is Anthropic's coding tool: you type a request in a terminal, the text window where you type commands, and an AI model reads files, runs commands and edits code. A token is a piece of text the model processes, and Anthropic's pricing page estimates one at 4 characters. The context window is all the text the model reads in one request, and Claude Code sends the whole conversation in every request. The prompt cache is a store of request text the service has already processed, which is billed at a lower price when it is read again.
Why measure at all
Three things Anthropic's documentation leaves open make a measurement worth doing. It gives no figure for what any single change saves. It gives no target for a good cache hit rate, the share of input read from the cache. And it states in general neither whether a request that misses the cache is billed as a cache write or as ordinary input, apart from three named cases. Our guide on how to reduce Claude Code token usage says each of these out loud, because a rule with a missing number is a rule a reader cannot budget on.
Reveneau is an AI software development consultancy. All of its code is written by AI, and every change must pass an eval suite, a set of automated tests written from the specification, before release. It uses AI instead of hiring more engineers, so a build takes a small team, and token use is a running cost of every Reveneau build. We want the numbers for our own work. Publishing them, with the method and the runs, is how a figure from a consultancy becomes something a reader can check instead of something a reader has to trust. Reveneau is independent of Anthropic.
The method, fixed before any run
- One public open-source repository at a pinned commit, chosen because it has a test suite. A pinned commit means every run starts from the same code.
- A fixed list of tasks, each with a test that passes or fails. A setting that saves tokens and fails the task is reported as a failure. A cheaper wrong answer is a wrong answer.
- Each setting is run several times, because the same request gives different results on different runs. The report gives the middle value and the full range, and publishes every run.
- Token counts come from Claude Code's own output. Anthropic's documentation says that with
--output-format jsonaclaude -prun, which is Claude Code started by a script with no person at the keyboard, reportstotal_cost_usdand a per-model breakdown. The counts are split into new input, output, cache reads and cache writes, because each is billed at a different rate: on Claude Sonnet 5.5, Anthropic's pricing page lists $2, $10, $0.20 and $2.50 per million tokens. Anthropic's Agent SDK page callstotal_cost_usda client-side estimate that can differ from the bill, so the token counts are the record and the dollars are arithmetic on list prices. - The Claude Code version and the model ID are pinned and printed on every table. Anthropic's model documentation says an alias such as
opusresolves to the recommended version and updates over time, and that a full name pins one version. Its pricing page says Claude 4.7 and later models use a newer tokenizer, the part that splits text into tokens, which Anthropic puts at 30 percent more tokens for the same text. A table without a version is a table about an unknown product. - One thing changes per experiment.
- The scripts, the task list and the raw results go in a public repository.
The eight experiments
Each one matches a page in the guide, so a reader can go from the advice to the measurement of it.
| Experiment | What changes | The page it tests |
|---|---|---|
| 1 | CLAUDE.md, the instruction file Claude reads at the start of every session, at three lengths | How long should CLAUDE.md be |
| 2 | Zero, three and ten tool servers connected | MCP servers or command-line tools |
| 3 | One long session against /clear between tasks |
/clear, /compact or /rewind |
| 4 | The same tasks on each current model | Which model and effort level |
| 5 | The same tasks at each effort level | The same page |
| 6 | Noisy test output read directly against output filtered by a hook, a script Claude Code runs by itself | Hooks that trim test and log output |
| 7 | Noisy work in the main session against the same work in a subagent, a second copy of Claude with its own context window | The same page |
| 8 | A message sent inside the cache lifetime against one sent after it has expired | How to read /usage and /context |
The pilot comes first
Step zero is a pilot: one task, two settings, five runs each. It gives two numbers the full benchmark cannot be planned without. The first is the cost per run, which sets what the whole benchmark will cost and the spending cap it runs under. The second is the spread between runs of the same setting, which sets how many runs each setting needs before a difference between two settings can be read as a difference and the gap between their ranges is wider than the gap inside one. No cost figure exists until the pilot has run.
What the benchmark cannot show
Two limits are fixed now, so that no result is read past them.
One repository is one repository. A result on it says how these settings behaved on that code, in that language, at that size, on those tasks. It says nothing about your repository. The scripts are published so that the same tasks can be run on other code, and a reader who runs them on their own repository learns more than our table can tell them.
A subscriber's plan-limit use cannot be read from token counts alone. Anthropic's pages, read on 4 October 2026, give no statement of how cache tokens count against the usage limits of a Pro, Max, Team or Enterprise plan. Every dollar figure in the report will therefore be an API list price, and a subscriber will see the token counts and have to judge the plan effect for themselves.
What we are promising
The report will live at the Claude Code token benchmark, with one page per experiment, its method, its table and a link to the raw files. The promise is the method above, the pinned versions, the pass-or-fail grading, several runs per setting, and every run public. The result and the date come when the runs are done. The method is published first so that it cannot be adjusted to fit a result afterwards.
A number with its working beside it is the only kind we will publish.
Sources
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026: the changes Anthropic names, and the absence of a saving figure.
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026: the cache, the full conversation sent on every request, and the three named cache-miss cases.
- Anthropic, Run Claude Code programmatically (code.claude.com), read 4 October 2026:
claude -p,--output-format jsonandtotal_cost_usd. - Anthropic, Track cost and usage, Agent SDK (code.claude.com), read 4 October 2026:
total_cost_usdas a client-side estimate. - Anthropic, Model configuration (code.claude.com), read 4 October 2026: aliases and pinned model names.
- Anthropic, Pricing (platform.claude.com), read 4 October 2026: the token estimate, the per-token prices, and the newer tokenizer.


