Method

How the Claude Code token benchmark is run

Reveneau's Claude Code token benchmark runs one small coding task through Claude Code's non-interactive mode (claude -p) five times per setting, in a fresh copy of the task folder each time, and records what Claude Code itself reports: the dollar figure, new input tokens, output tokens, cache read tokens, cache write tokens and the number of turns. The first dated run, on 4 October 2026, covered nine settings in four experiments, 45 recorded runs, all 45 passing, at $9.15 in total by Claude Code's own list-price figure on a subscription account. This page gives the whole method, so a reader can check each figure or repeat the run. The raw data is published.

Published October 4, 2026. Editorial.

Key takeaways

  • Every run used Claude Code 2.1.118 on one Mac, signed in with a Claude subscription, and the aliases resolved to claude-sonnet-4-6, claude-opus-4-7 and claude-haiku-4-5-20251001 on that day.
  • The task was a Node.js project of 30 lines across three source files, with one failing test out of four; after each run the test was re-run outside Claude Code, and exit 0 counted as a pass.
  • Each setting was run five times in a fresh copy of the folder, and the figures on every page of this guide are medians of those five runs.
  • The dollar figure is Claude Code's own total_cost_usd, which Anthropic's documentation, read on 4 October 2026, describes as a client-side estimate computed locally from a price table.
  • One run before the recorded ones was discarded because its test command named a directory that Node 22 rejected, so its pass or fail could not be read; the published README says so.

Reveneau is an AI software development consultancy. All of its code is written by AI, every change must pass an eval suite (a set of automated tests written from the specification) before release, and token use is therefore a running cost of every build. That is why Reveneau measures it. This page is the method behind the Claude Code token benchmark: what was run, how, and what was recorded, in enough detail that a reader can check every figure in the raw data or repeat the run on another machine. Reveneau is independent of Anthropic.

The short version: one small coding task, nine settings, five runs per setting, a fresh copy of the folder for every run, and only what Claude Code itself reports. The first dated run was on 4 October 2026. It produced 45 recorded runs, all 45 passed, and they cost $9.15 in total by Claude Code's own list-price figure on a subscription account. Every run is one row in the published runs.json.

The words used on every page of this guide

A token is a piece of text that the model reads or writes. Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters or 0.75 of an English word, and prices every model per million tokens [6].

The context window is all the text the model reads in one request: the instructions, the conversation so far, the files it has opened, the output of commands it has run, and the new message.

Claude Code sends that whole context again on every request. The service stores the unchanged start of a request so it does not have to process it again next time. Tokens stored this way are cache write tokens, and tokens read back from that store on a later request are cache read tokens. Anthropic's prompt caching page explains the mechanism and says a cache read is billed at a lower rate than new input [4]. The pricing page gives the rates [6].

CLAUDE.md is a file of instructions that Claude Code reads at the start of every session and includes in every request.

The effort level is a setting that tells the model how much reasoning to do before each answer.

A hook is a shell command that Claude Code runs on its own at a fixed point, for example before a command runs.

A non-interactive run is claude -p followed by a prompt: Claude Code does the task and exits, with no conversation in a terminal. Anthropic's documentation, read on 4 October 2026, describes this mode and the flags it accepts [1].

A turn, on these pages, is one request from Claude Code to the model and the reply that comes back. A run that reads a file, edits it and runs a test sends several requests, so one prompt produces several turns. Anthropic's cost tracking documentation calls one request and response cycle a step [2]. The benchmark's scripts read the count from the num_turns field of the json result.

A median is the middle value of five sorted runs: two runs are lower and two are higher.

The task and the prompt

The task is a Node.js project of three source files, src/money.js, src/ledger.js and src/report.js, 30 lines in total, and one test file with four tests. One test fails: the search() function must match the memo text without regard to upper and lower case, and it does not. The project is published under task/ beside the raw data.

The prompt for every run outside the hook experiment was:

One test in test/ledger.test.js fails. Fix the code in src/ so that node --test test/ledger.test.js passes. Change only src/. Do not change the tests. When done, stop.

The hook experiment's prompt named both test files, asked for one exact command to run them, and said that the second test prints a lot. The hook experiment page gives it in full.

The task was chosen to be small on purpose. It took a median of 6 to 10 turns per setting, so a whole setting of five runs cost under $3 by Claude Code's figure. The cost of that choice is that the run can say nothing about long sessions, which the limits page sets out.

The command and its flags

Every run called claude -p with the prompt and these flags. Anthropic's CLI reference, read on 4 October 2026, is the source for what each one does [3].

Flag What it does, by Anthropic's CLI reference Value used
-p Runs the prompt and prints the result instead of opening an interactive session [3] the prompt above
--model Sets the model for the session by alias or full name [3] sonnet, opus or haiku
--permission-mode Sets the permission mode the session starts in [3]; acceptEdits approves reads, file edits and common file-system commands without a prompt, by Anthropic's permission modes page [7], and other shell commands still need an allow rule [1] acceptEdits
--max-turns Limits the number of agentic turns (the reference's word for turns) and exits with an error at the limit; no limit by default [3] 25
--max-budget-usd Stops the run when the dollar estimate reaches the cap; applies in print mode, the -p mode, only [3] 2 in the CLAUDE.md experiment, 3 in the others
--output-format Chooses text, json or stream-json for print mode [3] json
--allowed-tools Names the tools that run without a permission prompt, using permission rule syntax [3] Read, Edit, Write, Glob, Grep, Bash(node:*), Bash(ls:*), Bash(cat:*)
--effort Sets the effort level for the session and does not persist [3] low, medium or high, in the effort experiment only
--settings Loads a settings file whose values override the same keys for this session [3] the hook settings file, in the hook experiment only

Standard input (the stream a program reads typed or piped text from) was redirected from /dev/null, an empty source, because a -p run reads standard input [1] and the benchmark wanted no input beyond the prompt. No run reached the turn cap or the dollar cap. The two caps were there as a safety limit for an unattended run, which is the practice the agent costs guide recommends.

Each run started in a fresh copy of the task folder. The script copied the folder, ran git checkout and git clean inside the copy so that it matched the committed state exactly, ran Claude Code there, graded the result, and deleted the copy. No run could see another run's edits. The published scripts, run.sh for the CLAUDE.md experiment and run2.sh for the others, contain these steps.

What was read from the json result

With --output-format json, Claude Code prints one JSON object (a block of structured text that a script can read) at the end of the run. Anthropic's documentation, read on 4 October 2026, says this payload includes total_cost_usd and a per-model cost breakdown, and that both are client-side estimates that can differ from an actual bill [1]. The benchmark read six values from it.

total_cost_usd is the dollar figure. Anthropic's cost tracking page says Claude Code computes it locally from a price table bundled at build time, at list price unless an administrator has set a modelPricing table, and that it can drift from a real bill when prices change or the installed version does not recognise a model [2]. The benchmark account is a Claude subscription, so there is no bill for these runs at all. The figure is the same quantity the /usage command shows in an interactive session, which Anthropic's costs page also describes as computed locally at list price [5]. On every page of this guide it is described as Claude Code's own list-price figure on a subscription account, because that is what it is.

usage.input_tokens is the count of new input tokens, the part of each request that was neither written to nor read from the cache. usage.output_tokens is the count the model wrote. Anthropic's cost tracking page says to read output tokens from the result message, because the per-step value is a placeholder [2]. The benchmark read all its counts from the result.

usage.cache_read_input_tokens and usage.cache_creation_input_tokens are the cache read and cache write counts. Anthropic's page says the first is charged at a reduced rate and the second at a higher rate than standard input [2].

num_turns is the turn count, and modelUsage names the model each request resolved to, which is how the resolved model names below were read.

Pass or fail

After each run, the command node --test test/ledger.test.js was run again, outside Claude Code, in the same copy of the folder. Exit code 0 (the number a command returns to say it succeeded) is a pass; anything else is a fail. The model's own statement that it had finished played no part. All 45 recorded runs passed, so every table in this guide compares settings that all did the job.

Five runs per setting, and one discarded run

Each setting was run five times. Five is enough to see the spread between identical runs beside the median, and small enough that nine settings cost $9.15 in total. On this task the spread was small: in seven of the nine settings the difference between the lowest and highest cost was $0.06 or less. The two exceptions were Opus, where one run took 159 seconds and cost $0.612 against four at $0.488 to $0.490, and the hook setting, at $0.164 to $0.245. Across all 45 runs the cost ranged from $0.060 (a Haiku run) to $0.612 (that Opus run), and wall-clock time (the time from start to finish on an ordinary clock) from 14 to 159 seconds.

The figures reported are medians of five, and each table gives the minimum and maximum beside the median. One task is one task: nothing in this run is a statistic about Claude Code in general.

One run before the recorded ones was discarded. Its test command named a directory, node --test test/, which Node 22 rejected, so the grading step could not read a pass or a fail. It is absent from the data, and the published README.txt says so.

The account, the cache lifetime, the version and the models

The machine was one Mac running Claude Code 2.1.118, signed in with a Claude subscription and no API key. Two consequences follow.

First, the dollar figures are the estimates Claude Code computes locally at list price, as described above. They are the comparable quantity across settings, and they are what /usage would show.

Second, every cache write was a one-hour write. Anthropic's prompt caching page, read on 4 October 2026, says Claude Code requests the one-hour cache lifetime for the main conversation only on a Claude subscription within plan usage, and five minutes on an API key or cloud provider; it also says the way to confirm the lifetime is to read usage.cache_creation in a json result, where one-hour writes appear under ephemeral_1h_input_tokens and five-minute writes under ephemeral_5m_input_tokens [4]. Reveneau checked the first run of each of the first two settings: all cache creation tokens were under the one-hour field and zero under the five-minute field. The pricing page lists a higher rate for one-hour writes than for five-minute writes [6], so a reader on an API key would see different cache write costs for the same runs. The prompt caching guide covers the two lifetimes.

The model aliases on that day and version resolved to claude-sonnet-4-6 for sonnet, claude-opus-4-7 for opus and claude-haiku-4-5-20251001 for haiku, read from the modelUsage field. These differ from the model names in Anthropic's current documentation, which names Opus 5.5, Sonnet 5.5 and Fable 5.1, because the installed Claude Code was 2.1.118. Prices and behaviour for the newer models were not measured.

Tokens per run were dominated by cache reads. In the no-CLAUDE.md Sonnet setting the median run had 184,772 cache read tokens, 25,579 cache write tokens, 8 new input tokens and 659 output tokens. Reveneau's reading of those counts: Claude Code re-sends the whole conversation on every request and the unchanged start is read from the cache, so a seven-turn run reads its growing unchanged start seven times. Anthropic's prompt caching page describes the same mechanism [4]. The token usage guide explains where each type of token comes from.

The nine settings

Setting Experiment Model as resolved Runs
No CLAUDE.md, default effort CLAUDE.md length; also the Sonnet row of the model experiment and the default row of the effort experiment claude-sonnet-4-6 5
301-line CLAUDE.md CLAUDE.md length claude-sonnet-4-6 5
Haiku, no CLAUDE.md Model claude-haiku-4-5-20251001 5
Opus, no CLAUDE.md Model claude-opus-4-7 5
--effort low Effort level claude-sonnet-4-6 5
--effort medium Effort level claude-sonnet-4-6 5
--effort high Effort level claude-sonnet-4-6 5
Second test that prints 3,000 lines, no hook Output-trimming hook claude-sonnet-4-6 5
Second test that prints 3,000 lines, through the hook Output-trimming hook claude-sonnet-4-6 5

Nine settings, five runs each, 45 runs. The first setting serves three experiments, which is why the experiment pages for model choice and effort level repeat its row.

Where the data is

Everything is at /benchmarks/claude-code-tokens/2026-10-04/: runs.json and runs.csv with one row per run, README.txt, the task files including the 301-line CLAUDE.md and the test that prints 3,000 lines, and the three scripts. A reader who wants to run the same method on their own task can follow the do-it-yourself page.

Common questions

What task did the Reveneau Claude Code token benchmark use?

The benchmark used one small Node.js project: three source files (src/money.js, src/ledger.js and src/report.js, 30 lines in total) and one test file with four tests, one of which fails because the search() function does not match memo text case-insensitively. The prompt asked Claude Code to fix the code in src/ so that node --test test/ledger.test.js passes, to change only src/, and to stop when done. The task folder is published with the raw data.

How many times was each setting in the benchmark run?

Each of the nine settings was run five times, for 45 recorded runs in total, all on 4 October 2026 with Claude Code 2.1.118. Five runs per setting were chosen so that the spread between identical runs could be seen beside the median. Every figure reported in this guide is the median of those five, and each results table also gives the lowest and highest cost in the five. One task is one task, and five runs do not make a statistic about Claude Code in general.

Why did the benchmark copy the task folder before every run?

The benchmark copied the task folder, then ran git checkout and git clean inside the copy, before every run, so that no run could see another run's edits. After the run the copy was deleted. Without this step, the second run of a setting would start with the bug already fixed by the first run and would report a smaller, meaningless cost. The published scripts run.sh and run2.sh contain these steps.

Which Claude Code flags did every benchmark run use?

Every run used --permission-mode acceptEdits, --max-turns 25, --max-budget-usd 2 for the CLAUDE.md experiment or 3 for the others, and --output-format json, with an allowed tool list of Read, Edit, Write, Glob, Grep and Bash limited to commands starting with node, ls or cat. Standard input was redirected from /dev/null. Anthropic's CLI reference, read on 4 October 2026, describes each flag, and no run reached the turn cap or the dollar cap.

Where does the dollar figure in the benchmark come from?

The dollar figure is the total_cost_usd field that Claude Code writes into its json result. Anthropic's documentation, read on 4 October 2026, says this field is a client-side estimate that the software computes locally from a price table at list price, and that it can differ from an actual bill. The benchmark ran on a Claude subscription, so no invoice exists for these runs. The figure is the same quantity the /usage command shows, which is why it was used to compare settings.

How was pass or fail decided in the benchmark?

After each run finished, the command node --test test/ledger.test.js was run again outside Claude Code, in the same copy of the folder. An exit code of 0 counted as a pass and anything else as a fail. All 45 recorded runs passed. The grading step was separate from Claude Code so that the model's own report of success played no part in the result; only the test's exit code did.

Why was one benchmark run discarded?

One run before the five recorded no-CLAUDE.md runs used a test command that named a directory, node --test test/, which Node 22 rejected, so the grading step could not read a pass or a fail. Reveneau discarded that run and it is absent from the published data. The published README.txt states this. No other run was discarded, and no recorded run reached the turn cap or the dollar cap.

Why do the benchmark figures use medians of five runs?

A median is the middle value when the five costs are sorted: two runs cost less and two cost more. It was chosen because one unusual run cannot move it. In the Opus setting one run took 159 seconds and cost $0.612 while the other four cost $0.488 to $0.490, and the median of $0.489 reports the usual run. Each table also gives the minimum and maximum, so the unusual run stays visible.

Which cache lifetime did the benchmark's cache writes use?

Every cache write in the benchmark was a one-hour write. Reveneau checked the json result of the first run in each of the first two settings: all cache_creation tokens were reported under ephemeral_1h_input_tokens and zero under ephemeral_5m_input_tokens. Anthropic's prompt caching page, read on 4 October 2026, says this is the default for the main conversation on a Claude subscription within plan usage, and names those two fields as the way to confirm it.

Which model versions did the benchmark measure?

The benchmark's model aliases resolved to claude-sonnet-4-6, claude-opus-4-7 and claude-haiku-4-5-20251001, read from the modelUsage field of each json result. These differ from the models named in Anthropic's current documentation, which names Opus 5.5, Sonnet 5.5 and Fable 5.1, because the installed Claude Code was version 2.1.118. Prices and behaviour for newer models were not measured in this run.

What does a turn mean in the benchmark tables?

On these pages a turn is one request from Claude Code to the model and the reply that comes back. A run that reads a file, edits it and runs a test sends several requests, one after each tool result, so a single prompt produces several turns. The benchmark's scripts read the turn count from the num_turns field of the json result. In the no-CLAUDE.md Sonnet setting the median was 7 turns.

Can I download the raw data from the Reveneau benchmark?

Yes. Every recorded run is published as one row in runs.json and runs.csv at /benchmarks/claude-code-tokens/2026-10-04/, with the cost, token counts, turn count, pass or fail and wall-clock seconds. The same folder holds README.txt, the task files including the 301-line CLAUDE.md and the test that prints 3,000 lines, and the scripts run.sh, run2.sh and filter-test-output.sh. The 45 rows there are the source of every figure in this guide.