Do it yourself

How to measure your own Claude Code runs the same way

To measure your own Claude Code runs the way Reveneau's benchmark of 4 October 2026 did, run claude -p with --output-format json and a turn cap and a dollar cap, start every run in a fresh copy of the project, grade each run by re-running a test outside Claude Code, run each setting five times, and report the median with the minimum and maximum. Record total_cost_usd, the four token counts in the usage object, the turn count and the resolved model name from each json result, along with the Claude Code version and whether the account is a subscription or an API key. This page gives each step, the fields and where they come from, and what the published scripts do.

Published October 4, 2026. Editorial.

Key takeaways

  • Run claude -p with --output-format json and read total_cost_usd, usage.input_tokens, usage.output_tokens, usage.cache_read_input_tokens and usage.cache_creation_input_tokens from the result; Anthropic's documentation, read on 4 October 2026, names each field.
  • Start every run in a fresh copy of the project, with git checkout and git clean inside the copy, so no run sees another run's edits.
  • Grade with a test that fails before the run and passes after it, re-run outside Claude Code, and keep only the exit code as the result.
  • Run each setting five times and report the median with the minimum and maximum; one unusual run cannot move a median, and the range keeps it visible.
  • Record the Claude Code version from claude --version, the resolved model from modelUsage, and whether the account is a subscription or an API key, because the cache lifetime and the meaning of the dollar figure differ between the two.

The figures in Reveneau's Claude Code token benchmark describe one task on one day. Your task is different, and the only figures that describe it are the ones you measure. This page is the method, step by step, so you can measure a setting on your own project the same way and compare two settings under the same conditions. Reveneau is an AI software development consultancy whose code is written by AI; token use is a running cost of every build, and this is how Reveneau measures it. Reveneau is independent of Anthropic.

The words used on this page

A token is a piece of text the model reads or writes. The context window is all the text the model reads in one request. Claude Code sends the whole conversation again on every request; the service stores the unchanged start, and tokens stored this way are cache write tokens, while tokens read back from the store on a later request are cache read tokens [6]. CLAUDE.md is a file of instructions loaded at the start of every session. The effort level is a setting for how much reasoning the model does before each answer. A hook is a shell command Claude Code runs on its own at a fixed point. A non-interactive run is claude -p with a prompt: Claude Code does the task and exits, and Anthropic's documentation, read on 4 October 2026, describes this mode [1]. A turn is one request to the model and its reply, so a run with several tool calls has several turns. A median is the middle value of your sorted runs.

Step 1: choose a task with a test

The task needs a check that does not depend on reading the model's answer. Reveneau's benchmark used a project with one failing test, and the prompt asked for the code to be fixed so the test passed. After each run the test was re-run outside Claude Code, and exit code 0 (the number a command returns to say it succeeded) was a pass. Choose a task of the same shape: something that fails before the run and passes after it, with a command you can run yourself to tell the difference.

Keep the task small while you are learning the method. Every Sonnet run of Reveneau's benchmark task, which took a median of 6 to 10 turns per setting, cost between $0.149 and $0.256 by Claude Code's list-price figure, so five runs per setting are affordable. A long task is a fair thing to measure later, once the method is working.

Step 2: write the prompt once and keep it fixed

Write the prompt in a file or a script and use the same text for every run of every setting. Reveneau's prompt named the test file, the command to run, the folder to change, and told the model to stop when done. If you change the prompt between settings, you have changed two things at once and the comparison is gone. The one exception in Reveneau's benchmark was the hook experiment, which needed a second test file in the command, and the hook experiment page says so.

Step 3: start every run in a fresh copy

Copy the project folder before every run. Inside the copy, run git checkout and git clean so the copy matches the committed state exactly, run Claude Code there, grade the result, and delete the copy. Without this, the second run starts with the first run's fix already in place, and its cost measures nothing. Reveneau's scripts do this with a temporary folder per run.

Step 4: the command and its flags

Run claude -p with your prompt and, at least, these flags. Anthropic's CLI reference, read on 4 October 2026, is the source for each [3].

--output-format json prints one JSON object (a block of structured text that a script can read) at the end instead of plain text [3]; this is where every figure comes from. --max-turns limits the number of agentic turns (the reference's word for turns) and exits with an error at the limit [3]. --max-budget-usd stops the run when the dollar estimate reaches the cap; it applies in print mode, the -p mode, only [3]. Set both caps above what a normal run needs, so that they stop a run that loops and never cut a normal run short; Reveneau used 25 turns and $2 or $3, and no run reached either. --permission-mode sets the mode the session starts in, and acceptEdits approves reads, file edits and common file-system commands without a prompt, by Anthropic's permission modes page [7]; --allowed-tools names the other tools and commands that may run without one, using permission rule syntax such as Bash(node:*) [3]. Redirect standard input (the stream a program reads typed or piped text from) from /dev/null, an empty source, because a -p run reads standard input [1].

Do not use --continue or --resume between runs. Anthropic's documentation says a continued or resumed run reports the conversation's whole total, earlier runs' spend included [1][2], and that a resumed conversation reads its earlier history from the cache [6]. Each run must be a new session so that its json result covers that run alone.

Anthropic recommends --bare for scripted calls, which skips loading hooks, skills, CLAUDE.md and the rest of the working directory's configuration, and which does not use a subscription login [1]. Reveneau did not use it, because the CLAUDE.md experiment needed CLAUDE.md to load. If your question is about a bare run, use the flag; if it is about what a developer's session costs, leave it off, and say which you did.

Step 5: grade, then read the result

After the run, re-run your test command yourself and record the exit code. Then read the json result. The table gives each field, what it counts, and where Anthropic documents it.

Field in the json result What it counts Source
total_cost_usd The dollar figure, computed locally from a bundled price table at list price unless a modelPricing table is in effect; a client-side estimate that can differ from a bill Anthropic, Track cost and usage [2]; Run Claude Code programmatically [1]
usage.input_tokens New input tokens, the part of each request neither written to nor read from the cache Anthropic, Track cost and usage [2]
usage.output_tokens Tokens the model wrote, including thinking; read it from the result, because per-step values are placeholders Anthropic, Track cost and usage [2]; Manage costs effectively [4]
usage.cache_read_input_tokens Tokens read from the cache, charged at a reduced rate Anthropic, Track cost and usage [2]
usage.cache_creation_input_tokens Tokens written to the cache, charged at a higher rate than standard input Anthropic, Track cost and usage [2]
usage.cache_creation.ephemeral_1h_input_tokens and ephemeral_5m_input_tokens The same cache writes split by lifetime, one hour or five minutes Anthropic, How Claude Code uses prompt caching [6]
modelUsage A map from each model's full name to its token counts and cost, which is how you learn what an alias resolved to Anthropic, Track cost and usage [2]
session_id The session's identifier Anthropic, Run Claude Code programmatically [1]
num_turns The turn count; Reveneau's published scripts read it from the result the published run.sh and run2.sh

Record all of these for every run, plus the wall-clock seconds (start to finish on an ordinary clock) from your own clock and the exit code of the test. One row per run, as in the published runs.json and runs.csv.

Step 6: five runs, then medians and ranges

Run each setting five times. Identical runs do not cost the same: in Reveneau's benchmark the five low-effort runs on Sonnet ranged from $0.149 to $0.208 by Claude Code's figure, and one of five Opus runs cost $0.612 against $0.488 to $0.490 for the other four. A single run of each setting would have reported whichever of those happened to occur.

Report the median of the five with the minimum and maximum beside it. The median is the middle value when the five are sorted, so one unusual run cannot move it, and the range shows that run instead of hiding it. Reveneau's tables give the median cost, its range, and the medians of the four token counts, the turn count and the seconds. Keep the raw rows and publish them if you publish the medians, so a reader can compute anything you did not.

Then say what the figures are. One task is one task. Medians of five runs describe that task on that day with that version, and nothing wider. Reveneau's limits page is the list for its own run; write yours.

Step 7: record the setup

Four facts belong beside the numbers, because each changes what they mean.

The Claude Code version, from claude --version. Anthropic's costs page says Claude Code regularly receives updates that may change how features work, including cost reporting [4]. Reveneau's run was on 2.1.118.

The resolved model names, from modelUsage. In Reveneau's run sonnet resolved to claude-sonnet-4-6, opus to claude-opus-4-7 and haiku to claude-haiku-4-5-20251001, and Anthropic's current documentation names newer models, so the alias alone is not enough.

Whether the account is a subscription or an API key. Anthropic's prompt caching page says the main conversation gets a one-hour cache lifetime on a Claude subscription within plan usage and five minutes on an API key or cloud provider by default, and that usage.cache_creation in the json result shows which applied [6]. The pricing of a one-hour write differs from a five-minute write, so cache write figures from the two setups are not comparable without this fact. Reveneau's benchmark ran on a subscription, and the dollar figures on every page of this guide are described as Claude Code's own list-price figure on a subscription account for that reason. The prompt caching guide covers the two lifetimes.

The date. Prices, aliases and defaults change, and a figure without a date cannot be checked.

The published scripts

Three scripts are published with the data, and a reader can run or edit them.

run.sh ran the CLAUDE.md experiment. It takes a setting letter, A or B, and a run number. It copies the task folder to a temporary directory, runs git checkout and git clean there, copies the 301-line CLAUDE.md into the folder for setting B only, runs claude -p with the benchmark flags on sonnet with a $2 cap, re-runs the test to grade, reads the fields above from the json result, appends one row to a results file, and deletes the temporary directory.

run2.sh ran the model, effort and hook experiments. It takes its choices as environment variables: the model, the effort level, whether to add the CLAUDE.md, whether to attach the hook with --settings, and whether to add the test that prints 3,000 lines, which also switches the prompt. It runs the same loop with a $3 cap and records the model names from modelUsage in each row.

filter-test-output.sh is the hook used in the fourth experiment, Anthropic's example with two strings changed, which the hook experiment page describes.

The interactive equivalents

In an interactive session the same figures appear in two places. The /usage command's Session block shows the total cost, computed locally at list price, and the token counts by model, which Anthropic's costs page describes [4]; the sister page on how to read /usage and /context explains each line. A status line script receives cost.total_cost_usd, which Anthropic's status line page describes as the estimated session cost computed client-side at list price, along with the context_window fields [5]; the sister page on checking your cache hit rate covers the cache fields it also receives. Both show one session as it runs. For a comparison between two settings you need five fresh runs of each, and that is what claude -p in a script gives you.

Common questions

How do I repeat the Reveneau benchmark method on my own task?

Pick a task with a test that fails before the run and passes after it. Run claude -p with your prompt, --output-format json, --max-turns and --max-budget-usd, in a fresh copy of the project each time. After each run, re-run the test outside Claude Code and record the exit code. Read the cost and token fields from the json result. Do five runs per setting and report medians with ranges. The published run2.sh does all of this and can be edited for your project.

Which json fields should I record from each claude -p run?

Record total_cost_usd, the four counts in the usage object (input_tokens, output_tokens, cache_read_input_tokens and cache_creation_input_tokens), the turn count, and the model names in modelUsage. Anthropic's documentation, read on 4 October 2026, describes total_cost_usd as a client-side estimate and says to read output tokens from the result instead of from per-step messages. If you want the cache lifetime, also read usage.cache_creation, which splits writes into one-hour and five-minute fields.

Why must each measured run start from a fresh copy of the project?

Because the first run changes the project. If the second run of a setting starts in the same folder, the bug is already fixed, the test already passes, and the run reports a smaller cost that measures nothing. Reveneau's scripts copy the project folder, run git checkout and git clean inside the copy so it matches the committed state, run Claude Code there, grade, and delete the copy. Five runs of one setting then each start from the same bytes.

How many runs per setting does Reveneau's method use, and why?

Five. Identical runs of Claude Code do not cost the same: in Reveneau's benchmark the five low-effort runs ranged from $0.149 to $0.208 by Claude Code's figure, and one of five Opus runs cost $0.612 against $0.488 to $0.490 for the others. One run of each setting would have reported whichever of those happened. Five is enough to see the spread beside the median, and nine settings of five cost $9.15 in total.

Should I report the mean or the median of my runs?

Report the median, with the minimum and maximum beside it. A median is the middle value of the sorted runs, so one unusual run cannot move it; a mean of five runs that include a $0.612 run beside four at $0.488 to $0.490 would be raised by that one run. The range shows the unusual run so it is not hidden. Reveneau's benchmark tables give exactly these three figures for each setting, and the raw rows let a reader compute anything else.

What should I record about my setup beside the token counts?

Four things. The Claude Code version, from claude --version, because Anthropic's costs page says updates can change how features work, including cost reporting. The model each alias resolved to, from modelUsage in the json result. Whether the account is a subscription or an API key, because Anthropic's prompt caching page gives the main conversation a one-hour cache lifetime on a subscription within plan usage and five minutes on an API key. And the date, because prices and aliases change.

How do I find out which model my alias resolved to?

Read the modelUsage field in the json result of a claude -p run. Anthropic's cost tracking page, read on 4 October 2026, describes it as a map from model name to that model's token counts and cost, so its keys are the full model names that served the run. In Reveneau's benchmark on 4 October 2026 with Claude Code 2.1.118, sonnet resolved to claude-sonnet-4-6, opus to claude-opus-4-7 and haiku to claude-haiku-4-5-20251001. Record the names, since an alias can resolve differently after an update.

Why does the benchmark grade with a test instead of reading Claude's answer?

Because the model's statement that it has finished is a claim, and the test is a check. Reveneau's method re-runs the test command outside Claude Code after the run and treats exit code 0 as a pass, so the pass column in every table rests on the code's behaviour alone. A run that reports success and leaves the test failing would be recorded as a fail. All 45 recorded runs in the 4 October 2026 benchmark passed by this check.

What do the published scripts run.sh and run2.sh do?

run.sh ran the CLAUDE.md experiment: it takes a setting letter and a run number, copies the task folder, adds the 301-line CLAUDE.md for setting B, runs claude -p with the benchmark flags, re-runs the test, and appends one row of results to a file. run2.sh ran the other three experiments and takes the model, effort level, CLAUDE.md, hook and second-test choices as environment variables, adds the hook with --settings when asked, and records the resolved model names from modelUsage.

Can I use --continue to run my five runs in one session?

No, if you want five independent figures. Anthropic's documentation, read on 4 October 2026, says that when you continue a conversation with --continue or --resume, the run reports the conversation's whole total, earlier runs' spend included, and that a resumed call also reads the earlier history from the cache. Each benchmark run was a new session in a new folder with no --continue, so each json result covers that run alone. Start a fresh session for every run.

How do I measure the same thing in an interactive session?

Run /usage, whose Session block shows the total cost computed locally at list price and the token counts by model, which Anthropic's costs page, read on 4 October 2026, describes as the same type of estimate as total_cost_usd. A status line script can read cost.total_cost_usd and the context_window fields on every turn. Neither gives you five identical runs in fresh folders, so for a comparison between settings the claude -p method is the one to use.

Do I need an API key to repeat the benchmark?

No. Reveneau's benchmark ran on a Claude subscription, and the dollar figure is Claude Code's local list-price estimate either way. Two things differ with an API key: Anthropic's prompt caching page says the main conversation's cache lifetime is five minutes by default instead of one hour, and the dollar figure then corresponds to a real bill at list price, which Anthropic's cost tracking page still calls an estimate. Record which one you used, so a reader can compare your cache write figures with the benchmark's.