Guide

The Claude Code token benchmark: one task, nine settings, 45 runs

Reveneau's Claude Code token benchmark ran the same small coding task through Claude Code's non-interactive mode (claude -p) 45 times on 4 October 2026, five runs per setting across nine settings in four experiments, and recorded what Claude Code itself reported for each run: the dollar figure, new input tokens, output tokens, cache read tokens, cache write tokens and the number of turns. All 45 runs passed their test, and they cost $9.15 in total by Claude Code's own list-price figure on a subscription account, with Claude Code 2.1.118 and aliases resolved to claude-sonnet-4-6, claude-opus-4-7 and claude-haiku-4-5-20251001. This guide is the report: the method, the four results with their arithmetic, what the run cannot show, and the raw data.

Published October 4, 2026. Editorial.

Key takeaways

  • A 301-line CLAUDE.md of rules the task did not need raised the median cost from $0.161 to $0.239, which is $0.078 or 48 percent by Reveneau's arithmetic, on claude-sonnet-4-6 over medians of five runs.
  • On the same task Haiku's median cost was 37 percent of Sonnet's and Opus's was 3.0 times Sonnet's, by Reveneau's arithmetic; Haiku took a median of 9 turns against 7 and still passed every run.
  • The effort level moved the median cost by $0.020 between low and high, 13 percent of the low median, and the ranges overlap; the difference came from one turn, with output token medians within 33 tokens of each other.
  • Anthropic's example test-output hook cost more on this task, a median of $0.201 against $0.159 and 10 turns against 7, because the filter removed the success summary and the model ran the tests again; the direct runs had already added tail -20 to the test command on their own, keeping only the last 20 lines of output.
  • One task is one task: the task took a median of 6 to 10 turns per setting, the figures are medians of five, the dollar figures are Claude Code's own list-price estimates on a subscription, and the models are the ones Claude Code 2.1.118 resolved to on that day.
  • Every run, the task files and the scripts are published at /benchmarks/claude-code-tokens/2026-10-04/, and a later dated run is planned with no schedule.

Reveneau ran the same small coding task through Claude Code's non-interactive mode, claude -p, 45 times on 4 October 2026: five runs per setting across nine settings in four experiments. For each run it recorded what Claude Code itself reported, the dollar figure (total_cost_usd, which Claude Code computes locally at list price), new input tokens, output tokens, cache read tokens, cache write tokens and the number of turns. After each run the test was re-run to grade the result pass or fail. All 45 runs passed. The runs cost $9.15 in total at Claude Code's list-price figure. Every run, the task files and the scripts are published at /benchmarks/claude-code-tokens/2026-10-04/runs.json.

Reveneau is an AI software development consultancy. All of its code is written by AI, every change must pass an eval suite (a set of automated tests written from the specification) before release, and it uses AI instead of hiring more engineers, so a build takes a small team. Token use is therefore a running cost of every Reveneau build, which is why Reveneau measures it. Reveneau is independent of Anthropic: the benchmark is Reveneau's own measurement of Anthropic's product, and Anthropic's documentation is cited on every page for what each setting does.

This page is the report. The method, each experiment, the do-it-yourself steps and the limits each have their own page, linked below.

The words used in this guide

A token is a piece of text that the model reads or writes. Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters or 0.75 of an English word, and prices every model per million tokens [5]. The context window is all the text the model reads in one request. Claude Code sends that whole context again on every request; the service stores the unchanged start of a request, and tokens stored this way are cache write tokens, while tokens read back from that store on a later request are cache read tokens, billed at a lower rate than new input by Anthropic's prompt caching page [2]. CLAUDE.md is a file of instructions that Claude Code reads at the start of every session and includes in every request. The effort level is a setting that tells the model how much reasoning to do before each answer. A hook is a shell command Claude Code runs on its own at a fixed point, for example before a command runs [7]. A non-interactive run is claude -p followed by a prompt: Claude Code does the task and exits [3]. A turn, on these pages, is one request from Claude Code to the model and its reply, so a run with several tool calls has several turns. A median is the middle value of five sorted runs.

The setup in short

One Mac, Claude Code 2.1.118, signed in with a Claude subscription. The dollar figures are the estimates Claude Code computes locally at list price from the token counts; Anthropic's cost tracking page, read on 4 October 2026, calls total_cost_usd a client-side estimate that can differ from a bill [4], and on a subscription there is no bill at all. They are the figures /usage would show [1] and the comparable quantity across settings. Because the account was within its plan usage, every cache write was a one-hour write, which Anthropic's prompt caching page says is the default for the main conversation on a subscription [2]; Reveneau confirmed it in the json result of the first run of each of the first two settings.

The model aliases on that day resolved to claude-sonnet-4-6, claude-opus-4-7 and claude-haiku-4-5-20251001. These differ from the model names in Anthropic's current documentation, which names Opus 5.5, Sonnet 5.5 and Fable 5.1 [6], because the installed Claude Code was 2.1.118.

Every run used --permission-mode acceptEdits, --max-turns 25, --max-budget-usd 2 or 3, --output-format json, and an allowed tool list of Read, Edit, Write, Glob, Grep and Bash limited to commands starting with node, ls or cat [3]. No run reached either cap. Each run started in a fresh copy of the task folder, so no run could see another run's edits. The task was a Node.js project of three source files, 30 lines in total, and one test file with four tests, one of which fails. Pass or fail was read by re-running the test outside Claude Code. One earlier run was discarded because its test command named a directory that Node 22 rejected. The method page gives all of this in full, with the prompt and every flag.

All nine settings in one table

Every figure is Reveneau's own measurement from the run of 4 October 2026. Cost is Claude Code's own list-price figure on a subscription account, the median of five runs, with the lowest and highest of the five beside it.

Setting Model as resolved Pass Median cost Lowest to highest Median turns
No CLAUDE.md, default effort claude-sonnet-4-6 5 of 5 $0.161 $0.161 to $0.163 7
301-line CLAUDE.md claude-sonnet-4-6 5 of 5 $0.239 $0.221 to $0.256 8
Haiku, no CLAUDE.md claude-haiku-4-5-20251001 5 of 5 $0.060 $0.060 to $0.071 9
Opus, no CLAUDE.md claude-opus-4-7 5 of 5 $0.489 $0.488 to $0.612 7
--effort low claude-sonnet-4-6 5 of 5 $0.150 $0.149 to $0.208 6
--effort medium claude-sonnet-4-6 5 of 5 $0.162 $0.149 to $0.201 7
--effort high claude-sonnet-4-6 5 of 5 $0.170 $0.161 to $0.174 7
Test that prints 3,000 lines, no hook claude-sonnet-4-6 5 of 5 $0.159 $0.149 to $0.173 7
Test that prints 3,000 lines, through the hook claude-sonnet-4-6 5 of 5 $0.201 $0.164 to $0.245 10

The first row serves three experiments: it is the no-CLAUDE.md setting, the Sonnet row of the model experiment and the default row of the effort experiment. Across all 45 runs the cost per run ranged from $0.060 to $0.612 and the wall-clock time (the time from start to finish on an ordinary clock) from 14 to 159 seconds. In seven of the nine settings the difference between the lowest and highest of five identical runs was $0.06 or less; the exceptions were Opus, where one run took 159 seconds, and the hook setting.

Finding 1: a 301-line CLAUDE.md the task did not need

With a CLAUDE.md of 301 lines of rules unrelated to the task (342 lines counting the blank lines between its sections, 62,746 characters), the median cost on claude-sonnet-4-6 rose from $0.161 to $0.239. By Reveneau's arithmetic that is $0.078, which is 48 percent of the no-CLAUDE.md median. Cache write tokens rose by 14,487 per run, which fits the file being written to the cache once per run, and cache read tokens rose by 71,678, which fits it being read back on every later request of a run with a median of 8 turns. The task needed nothing in the file. Anthropic's documentation, read on 4 October 2026, says to keep CLAUDE.md under 200 lines [1]; the experiment measured 301 lines against none and tested nothing in between. The CLAUDE.md length experiment has the full table, and the sister page on how long a CLAUDE.md should be covers Anthropic's guidance.

Finding 2: Haiku, Sonnet and Opus on the same task

With no CLAUDE.md and default effort, the median cost was $0.060 on claude-haiku-4-5-20251001, $0.161 on claude-sonnet-4-6 and $0.489 on claude-opus-4-7. By Reveneau's arithmetic, Haiku's median was 37 percent of Sonnet's and Opus's was 3.0 times Sonnet's. Haiku took more turns, a median of 9 against 7, and still passed every run. Opus read more cache tokens per run, a median of 367,005 against 184,772 for Sonnet with the same median of 7 turns; the unchanged start of each request, which it re-read every time, was larger; Reveneau did not verify why, and the pages do not guess. One Opus run took 159 seconds against 18 to 22 for the other four, with 8 turns against 7 for the other four, and the cause was not recorded. Anthropic's list prices for the three models are on the model choice experiment page with the date they were read, and Anthropic's costs page advises Sonnet for most coding tasks and Opus for complex architectural decisions or multi-step reasoning [1].

Finding 3: low, medium and high effort on Sonnet

Passed with --effort on claude-sonnet-4-6, the median cost was $0.150 at low, $0.162 at medium and $0.170 at high. By Reveneau's arithmetic the difference between low and high is $0.020, which is 13 percent of the low median, and the ranges overlap: one low run cost $0.208, above every high run. The difference in medians came from turns, 6 at low against 7 at the other two, which changes how many times the growing conversation is re-read from the cache. The output token medians were 644, 677 and 655, within 33 tokens of each other. Anthropic's documentation says thinking tokens are billed as output tokens [1], and on a task this small the model did not reason at length at any level, so the experiment cannot show what effort does on a task that needs it. The effort level experiment has the table, including the default setting with no flag, which Reveneau reports separately because it did not verify which level version 2.1.118 applied.

Finding 4: Anthropic's example test-output hook

A second test that prints 3,000 lines was added, and the task was run with and without Anthropic's example PreToolUse hook from its costs page, with two strings changed: the test-runner pattern set to match node --test, and not ok added to the words the filter keeps [1][7]. The hook runs cost more: a median of $0.201 against $0.159, $0.042 more by Reveneau's arithmetic, with a median of 10 turns against 7. Reveneau read all ten transcripts. In every direct run the model itself appended | tail -20 to the test command, which keeps only the last 20 lines of output, so the 3,000 lines never entered the conversation and the largest tool result was 869 characters. In the hook runs the filtered command produced no output on a pass, because the pattern matches only failure words and Node's summary lines contain none; the model saw only Claude Code's 31-character notice that the command had completed with no output, so it ran the tests again in other forms, between 2 and 8 test commands per run against 1, and one run reached 15 turns. Reveneau's reading: a filter that removes the success summary costs turns, and a filter should keep the summary line. Anthropic's costs page illustrates the hook with a 10,000-line log file; that is Anthropic's illustration, and this experiment did not test it. The output-trimming hook experiment has the transcript findings and the published script.

What the run cannot show, in short

One invented task of one failing test in a 30-line project is one task, and the medians are medians of five. Nothing here is a statistic about Claude Code in general. The task took a median of 6 to 10 turns per setting, so the run cannot show what a long session costs, what /clear (which empties the conversation) or /compact (which replaces it with a summary) saves, or what happens when the cache expires, which Anthropic's costs page lists among the reasons a session open for hours uses more of a plan than the activity suggests [1]. The dollar figures are Claude Code's local list-price estimates on a subscription and do not show how much of a plan's five-hour or weekly allowance a run used, which no Claude Code output reports. The models are the ones 2.1.118 resolved to; prices and behaviour for Opus 5.5, Sonnet 5.5 and Fable 5.1 were not measured, and Anthropic's pricing page says Claude 4.7 and later models count the same text as more tokens [5]. The CLAUDE.md experiment measured a file the task did not need. Four planned experiments were not run: tool servers (MCP servers, programs that give Claude Code extra tools), one long session against /clear between tasks, a subagent (a separate conversation that the main one starts for part of the work) against the main session, and inside against after the cache lifetime. The limits page gives each of these its own section and a table of claims the run does and does not support.

Where the raw data is

Everything is at /benchmarks/claude-code-tokens/2026-10-04/: runs.json and runs.csv with one row per run, README.txt, the task files including the 301-line CLAUDE.md and the test that prints 3,000 lines, and the scripts run.sh, run2.sh and filter-test-output.sh. Every median in this guide can be recomputed from the 45 rows, and a reader who wants a different summary can compute it.

How to measure your own runs

The figures above describe one task. The do-it-yourself page gives the method as steps: a task with a test, a fixed prompt, a fresh copy per run, the claude -p flags with a turn cap and a dollar cap, the json fields to record and where Anthropic documents each, five runs, medians with ranges, and the four facts to record beside the numbers: the version, the resolved model, whether the account is a subscription or an API key, and the date.

What comes next

A later dated run is planned, with no schedule. It is intended to reach the four experiments from the plan that this run did not, and each new run will have its own date in its path and state its version and resolved models, so that two runs can be compared without assuming the product stayed the same between them. Anthropic's costs page, read on 4 October 2026, says Claude Code regularly receives updates that may change how features work, including cost reporting [1], which is why the version is stated beside every figure. The first run's data stays published as it is.

Where this guide fits with the other Claude Code guides

This guide reports measurements. Anthropic's advice on each setting, and the reasons behind it, are covered in the guides published beside it. Reducing Claude Code token usage covers the changes that lower usage, including the three how-to pages that match the experiments here: CLAUDE.md length, model and effort, and hooks that trim output. Claude Code prompt caching explains the cache reads and writes that make up most of every run in the tables above, and what resets them. Claude Code costs for teams covers Anthropic's published per-developer averages, which are a different type of number from a median of five runs, and where a team sees and caps spend. Claude Code agent costs covers subagents, workflows and unattended runs with a spend cap, which is how the benchmark's own runs were capped. Claude Code usage limits covers the five-hour and weekly windows that a subscription run draws on and that no figure in this benchmark can measure.

Explore the guide

Experiments

The CLAUDE.md length experiment: 301 lines against none

In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times with no CLAUDE.md and five times with a 301-line CLAUDE.md of project rules unrelated to the task (62,746 characters; 342 lines when the blank lines between sections are counted), on claude-sonnet-4-6 at default effort with Claude Code 2.1.118. The median cost rose from $0.161 to $0.239 by Claude Code's own list-price figure on a subscription account, which is $0.078 or 48 percent more by Reveneau's arithmetic. Cache write tokens rose by 14,487 and cache read tokens by 71,678. All ten runs passed. One task is one task, the figures are medians of five, and the experiment measured one file against none, with nothing in between.

The model choice experiment: Haiku, Sonnet and Opus on one task

In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times on each of three models, with no CLAUDE.md and default effort, using Claude Code 2.1.118 on a Claude subscription. The aliases resolved to claude-haiku-4-5-20251001, claude-sonnet-4-6 and claude-opus-4-7. The median costs by Claude Code's own list-price figure were $0.060, $0.161 and $0.489. By Reveneau's arithmetic, Haiku's median was 37 percent of Sonnet's and Opus's was 3.0 times Sonnet's. All 15 runs passed. Haiku took more turns, a median of 9 against 7. One task is one task, the figures are medians of five, and newer models were not measured.

The effort level experiment: low, medium and high on Sonnet 4.6

In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times at each of three effort levels, passed with the --effort flag, on claude-sonnet-4-6 with no CLAUDE.md and Claude Code 2.1.118 on a Claude subscription. The median costs by Claude Code's own list-price figure were $0.150 at low, $0.162 at medium and $0.170 at high. By Reveneau's arithmetic the gap between low and high is $0.020, or 13 percent of the low median, and the ranges overlap: one low run cost $0.208, above every high run. The difference in medians came from turns, 6 against 7; output token medians are within 33 tokens of each other. All 15 runs passed. One task is one task, and the figures are medians of five.

The output-trimming hook experiment: a test that prints 3,000 lines

In Reveneau's Claude Code token benchmark of 4 October 2026, a second test printing 3,000 lines was added and the task ran five times with no hook and five times through Anthropic's example PreToolUse hook that filters test output, on claude-sonnet-4-6 with Claude Code 2.1.118 on a Claude subscription. The hook runs cost more: a median of $0.201 against $0.159 by Claude Code's own list-price figure, with a median of 10 turns against 7. The transcripts show why: in every direct run the model itself added | tail -20 to the test command, keeping the last 20 lines, and in the hook runs the filter left no output on a pass, so the model ran the tests again. One task is one task; the figures are medians of five.

Common questions

What is the Reveneau Claude Code token benchmark?

It is a measured run of one small coding task through Claude Code's non-interactive mode, claude -p, 45 times on 4 October 2026: five runs per setting across nine settings in four experiments, on CLAUDE.md length, model, effort level and an output-trimming hook. Reveneau recorded what Claude Code itself reported for each run, graded each run by re-running the test, and published every row, the task and the scripts. Every figure in this guide comes from those 45 rows.

What did the Reveneau benchmark find?

Four things, each about one task on 4 October 2026 with Claude Code 2.1.118. A 301-line unused CLAUDE.md raised the median cost by 48 percent on claude-sonnet-4-6. Haiku's median cost was 37 percent of Sonnet's and Opus's was 3.0 times Sonnet's. The effort level moved the median by 13 percent between low and high, with overlapping ranges. Anthropic's example test-output hook cost more, because its filter removed the success summary. All percentages are Reveneau's arithmetic on medians of five.

How much did the whole benchmark cost to run?

The 45 recorded runs cost $9.15 in total by Claude Code's own list-price figure, which is the total_cost_usd field of each json result added up. Cost per run ranged from $0.060, a Haiku run, to $0.612, an Opus run that took 159 seconds. The account was a Claude subscription within its plan usage, so the figure is Claude Code's local estimate at list price and no invoice exists for it. Anthropic's documentation, read on 4 October 2026, calls this field a client-side estimate.

Which setting cost the least and which the most in the benchmark?

By median cost, Haiku with no CLAUDE.md cost the least at $0.060 and Opus with no CLAUDE.md the most at $0.489, by Claude Code's own list-price figure on a subscription account, medians of five runs each. Among the Sonnet settings, low effort cost the least at $0.150 and the 301-line CLAUDE.md the most at $0.239. The single lowest-cost run was a Haiku run at $0.060 and the single most costly an Opus run at $0.612. All 45 runs passed.

Why does Reveneau run a token benchmark at all?

Because Reveneau is an AI software development consultancy whose code is written by AI, so tokens are a running cost of every build, and a setting that changes the cost of every request is worth measuring before it is adopted. Anthropic's documentation gives advice on CLAUDE.md length, model choice, effort and hooks; the benchmark checks what each costs on one task, with the working shown, and publishes the result so a reader can check or repeat it.

Why does every benchmark page repeat the Claude Code version and the model names?

Because the figures belong to that version and those models. The installed Claude Code was 2.1.118, and on 4 October 2026 its aliases resolved to claude-sonnet-4-6, claude-opus-4-7 and claude-haiku-4-5-20251001, while Anthropic's current documentation names Opus 5.5, Sonnet 5.5 and Fable 5.1 with different prices and defaults. Anthropic's costs page also says Claude Code regularly receives updates that may change how features work, including cost reporting. A figure quoted without its version and model would be a claim about a product that may have changed.

Where is the raw data for the benchmark?

At /benchmarks/claude-code-tokens/2026-10-04/. The folder holds runs.json and runs.csv with one row per run, giving the cost, the four token counts, the turn count, pass or fail and wall-clock seconds; README.txt; the task files, including the 301-line CLAUDE.md and the test that prints 3,000 lines; and the scripts run.sh, run2.sh and filter-test-output.sh. Every median in this guide can be recomputed from those 45 rows.

How should the benchmark figures be read alongside Anthropic's own cost figures?

As different kinds of number. Anthropic's costs page, read on 4 October 2026, states an average of $13 per developer per active day across enterprise deployments, which is Anthropic's statement about its own customers over whole working days. The benchmark's figures are medians of five runs of one task whose median run took 21 seconds or less in every setting, measured by Reveneau. Neither can be derived from the other, and the team costs guide covers how to use Anthropic's averages.

What is the single largest effect the benchmark measured?

The model. Opus's median cost was 3.0 times Sonnet's on the same task, and Haiku's was 37 percent of Sonnet's, by Reveneau's arithmetic on medians of five. The next largest was the 301-line CLAUDE.md at 48 percent above the no-CLAUDE.md median. The effort level moved the median by 13 percent between low and high, with overlapping ranges, and the hook raised cost by $0.042 at the median. All four apply to one task on 4 October 2026 with Claude Code 2.1.118.

Did the benchmark cost anything beyond the Claude subscription?

No. The account was a Claude subscription within its plan usage, which is also why every cache write was a one-hour write, as Anthropic's prompt caching page, read on 4 October 2026, says is the default for the main conversation on a subscription. The $9.15 total is Claude Code's estimate, at Anthropic's list prices, of the tokens the 45 runs used, and the raw data states this beside every figure.

Will the benchmark be repeated?

A later dated run is planned, with no schedule. It is intended to cover the four experiments from the plan that this run did not reach: tool servers, one long session against /clear between tasks, a subagent against the main session, and inside against after the cache lifetime. Each new run will have its own date in its path and state its Claude Code version and resolved model names, so runs can be compared without assuming the product stayed the same.

Does the benchmark recommend any setting?

Only where Anthropic's documentation already advises it and the measurement agrees, and then as Reveneau's recommendation for tasks like this one: keep CLAUDE.md under Anthropic's 200 lines, start a task on Sonnet and switch to a larger model when the task needs it, set the effort level at the start of a session, and make any output filter keep the test runner's summary line. The benchmark does not recommend beyond this task, and a reader's own task is measured with the do-it-yourself page.