The model choice experiment: Haiku, Sonnet and Opus on one task
In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times on each of three models, with no CLAUDE.md and default effort, using Claude Code 2.1.118 on a Claude subscription. The aliases resolved to claude-haiku-4-5-20251001, claude-sonnet-4-6 and claude-opus-4-7. The median costs by Claude Code's own list-price figure were $0.060, $0.161 and $0.489. By Reveneau's arithmetic, Haiku's median was 37 percent of Sonnet's and Opus's was 3.0 times Sonnet's. All 15 runs passed. Haiku took more turns, a median of 9 against 7. One task is one task, the figures are medians of five, and newer models were not measured.
Published October 4, 2026. Editorial.
Key takeaways
- On the benchmark task the median cost was $0.060 on claude-haiku-4-5-20251001, $0.161 on claude-sonnet-4-6 and $0.489 on claude-opus-4-7, by Claude Code's own list-price figure on a subscription account, medians of five runs each.
- By Reveneau's arithmetic, Haiku's median cost was 37 percent of Sonnet's and Opus's was 3.0 times Sonnet's; all 15 runs passed the test.
- Haiku took a median of 9 turns against 7 for Sonnet and Opus, and still passed every run.
- Opus read a median of 367,005 cache tokens per run against 184,772 for Sonnet; Reveneau did not verify why and this page offers no reason.
- One Opus run took 159 seconds and 8 turns against 18 to 22 seconds and 7 turns for the other four, and cost $0.612; the cause was not recorded.
This is the second experiment in Reveneau's Claude Code token benchmark. It asks what the same small task costs on each of the three models that the aliases haiku, sonnet and opus resolved to on 4 October 2026. Reveneau is an AI software development consultancy whose code is written by AI, so the choice of model is a cost decision in every build, and Reveneau is independent of Anthropic. The full method is on the method page.
The words used on this page
A token is a piece of text the model reads or writes; Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters of English and prices every model per million tokens [3]. The context window is all the text the model reads in one request. Claude Code sends the whole conversation again on every request; the service stores the unchanged start, and tokens stored this way are cache write tokens, while tokens read back from the store on a later request are cache read tokens [4]. CLAUDE.md is a file of instructions loaded at the start of every session; this experiment used none. The effort level is a setting for how much reasoning the model does before each answer; this experiment passed no flag for it. A hook is a shell command Claude Code runs on its own at a fixed point; none was used. A non-interactive run is claude -p with a prompt: Claude Code does the task and exits. A turn is one request to the model and its reply. A median is the middle value of five sorted runs.
What was compared
Three settings, each run five times: --model haiku, --model sonnet and --model opus. Anthropic's CLI reference, read on 4 October 2026, says the --model flag takes an alias such as these or a model's full name [6]. On that day, in Claude Code 2.1.118, the aliases resolved to claude-haiku-4-5-20251001, claude-sonnet-4-6 and claude-opus-4-7, read from the modelUsage field of each json result, which Anthropic's cost tracking page describes as a map from model name to that model's token counts and cost [5].
Everything else was held constant: the task of one failing test in a 30-line Node.js project, the same prompt, no CLAUDE.md, no --effort flag, the flags on the method page, a Claude subscription, and a fresh copy of the folder for every run. The Sonnet setting is the same five runs as setting A of the CLAUDE.md experiment.
Because no --effort flag was passed, each model ran at whatever default effort level Claude Code 2.1.118 applied to it. Anthropic's model configuration page, read on 4 October 2026, states a default per model for the current version [2]; Reveneau did not verify that those defaults applied in 2.1.118, so this page does not state the level each model ran at.
The result
Every figure is Reveneau's own measurement from the run of 4 October 2026, a median of five, with the dollar figure being Claude Code's own list-price figure on a subscription account. The rows are in runs.json.
| Model (as resolved) | Pass | Cost median (min to max) | Cache read median | Cache write median | Output median | Turns median | Seconds median |
|---|---|---|---|---|---|---|---|
| claude-haiku-4-5-20251001 | 5 of 5 | $0.060 ($0.060 to $0.071) | 222,836 | 26,543 | 903 | 9 | 15 |
| claude-sonnet-4-6 | 5 of 5 | $0.161 ($0.161 to $0.163) | 184,772 | 25,579 | 659 | 7 | 21 |
| claude-opus-4-7 | 5 of 5 | $0.489 ($0.488 to $0.612) | 367,005 | 45,364 | 873 | 7 | 21 |
All 15 runs passed.
Reveneau's arithmetic
Haiku's median cost of $0.060 is 37 percent of Sonnet's $0.161. Opus's median of $0.489 is 3.0 times Sonnet's. Both are Reveneau's arithmetic on medians of five runs.
Haiku took more turns, a median of 9 against 7 for the other two, and wrote the most output tokens, a median of 903. It still cost the least per run and passed every run. On a task this small, two extra turns on the lowest-priced model cost less than the same work on a higher-priced one.
Opus read more cache tokens per run than Sonnet, a median of 367,005 against 184,772, with the same median of 7 turns, and wrote more to the cache, 45,364 against 25,579. Reveneau's reading goes no further than the counts: the unchanged start of each request, which the Opus runs re-read every time, was larger. Reveneau did not verify why, and this page does not guess. Finding out would need a comparison of the Opus and Sonnet transcripts, which was not done in this run.
One Opus run took 159 seconds of wall-clock time (the time from start to finish on an ordinary clock) and 8 turns, against 18 to 22 seconds and 7 turns for the other four, and cost $0.612 against $0.488 to $0.490. The cause was not recorded. That run is the maximum in the table and the most costly of all 45 runs; the median of $0.489 reports the usual Opus run.
Anthropic's list prices for the three models
Anthropic's pricing page, read on 4 October 2026, gives these list prices per million tokens [3]. Claude Code's total_cost_usd is computed from a bundled price table at list price, by Anthropic's cost tracking page [5], so these are the rates behind the dollar figures above. Every cache write in the benchmark was a one-hour write, as the method page explains, so the one-hour column is the one that applied.
| Model | New input | One-hour cache write | Cache read | Output |
|---|---|---|---|---|
| Claude Haiku 4.5 | $1 | $2 | $0.10 | $5 |
| Claude Sonnet 4.6 | $3 | $6 | $0.30 | $15 |
| Claude Opus 4.7 | $5 | $10 | $0.50 | $25 |
These are Anthropic's own prices for its own models. The measured cost ratio between Opus and Sonnet, 3.0, is larger than the ratio of any one pair of prices in the table by Reveneau's arithmetic, and the result table shows that the Opus runs also used more tokens of every kind. Both things are true at once; which part of the difference comes from price and which from token counts was not separated in this run.
The pricing page also says that Claude 4.7 and later models use a newer tokenizer (the rule a model uses to split text into tokens), and the limits page notes what that means for carrying these token counts to other models.
What Anthropic advises on model choice
Anthropic's costs page, read on 4 October 2026, says that Sonnet handles most coding tasks well and costs less than Opus, that Opus should be reserved for complex architectural decisions or multi-step reasoning, and that for simple subagent tasks (a subagent is a separate conversation the main one starts for part of the work) model: haiku can be set in the subagent configuration [1]. The same page names Opus left as the default model as one of the two usual causes of unexpectedly high spend on an API or cloud-provider plan [1].
Reveneau's measurement agrees with the direction of that advice on this task: the task is a one-line fix with a test to confirm it, Sonnet's median was $0.161 against Opus's $0.489, and Haiku's was $0.060. Reveneau's recommendation, drawn from Anthropic's documentation and consistent with this measurement, is to start a task on Sonnet and switch to a larger model only when the task needs it. The sister page on which model and effort level to choose covers Anthropic's guidance in full, and the agent costs guide covers giving a smaller model to a subagent.
That recommendation is for this task and tasks like it. A harder task could change the number of turns each model needs, and a model that takes more turns reads its growing conversation from the cache more times, as the token usage guide explains. Nothing in five runs of a one-line fix says what the three models cost on a task that needs reasoning.
Switching model mid-task was not measured
Anthropic's prompt caching page, read on 4 October 2026, says each model has its own cache, so switching models recomputes the entire request even when the content is identical [4]. Every benchmark run set its model at launch and never changed it, so no run paid that cost. The cost of a mid-task switch is covered in the prompt caching guide.
What this experiment cannot show
One invented task of one failing test in a 30-line project is one task. The medians are medians of five, and nothing here is a statistic about Claude Code in general.
The models were the ones Claude Code 2.1.118 resolved to on 4 October 2026. Anthropic's current documentation names Opus 5.5, Sonnet 5.5 and Fable 5.1, with different prices and defaults, and prices and behaviour for those models were not measured. The ratios on this page are ratios between claude-haiku-4-5-20251001, claude-sonnet-4-6 and claude-opus-4-7 and apply to no other pair.
The dollar figures are Claude Code's own list-price estimates on a subscription account. They do not show how much of a plan's allowance each run used, which no Claude Code output reports, and a plan may count models differently from the list prices; the usage limits guide covers what counts against a plan.
The task took a median of 6 to 10 turns per setting. A long session, a task that needs reasoning, or a task where one model fails and another passes would each give a different comparison, and the run measured none of them. A later dated run is planned, with no schedule.
Common questions
Which three models did the Reveneau benchmark compare?
The benchmark compared the three models that the aliases haiku, sonnet and opus resolved to in Claude Code 2.1.118 on 4 October 2026: claude-haiku-4-5-20251001, claude-sonnet-4-6 and claude-opus-4-7, read from the modelUsage field of each json result. Each ran the same task five times with no CLAUDE.md and no --effort flag, in a fresh copy of the folder, on a Claude subscription. The 15 rows are in the published runs.json.
What share of Sonnet's median cost did Haiku's median reach in the benchmark?
Haiku's median cost of $0.060 was 37 percent of Sonnet's median of $0.161, by Reveneau's arithmetic on medians of five runs, with the dollar figure being Claude Code's own list-price figure on a subscription account. The five Haiku runs ranged from $0.060 to $0.071. Haiku took a median of 9 turns against Sonnet's 7 and wrote a median of 903 output tokens against 659, and still cost less per run on this task.
By how much did the median Opus run exceed the median Sonnet run in the benchmark?
The median Opus run cost $0.489 against $0.161 for Sonnet, which is 3.0 times Sonnet's median by Reveneau's arithmetic, using Claude Code's own list-price figure on a subscription account. Four of the five Opus runs cost $0.488 to $0.490 and one cost $0.612. The Opus median run also read 367,005 cache tokens against 184,772 for Sonnet, with the same median of 7 turns. One task is one task, and these are medians of five.
Did Haiku pass the benchmark task every time?
Yes. All five Haiku runs passed, graded by re-running node --test test/ledger.test.js outside Claude Code and treating exit code 0 as a pass. Haiku took more turns to get there, a median of 9 against 7 for Sonnet and Opus, and wrote the most output tokens of the three, a median of 903. On a task of one failing test in a 30-line project, the extra turns did not stop it from passing or from costing the least.
Why did Opus read more cache tokens than Sonnet in the benchmark?
Reveneau did not verify why, and this page gives no reason. The measurement is that the median Opus run read 367,005 cache tokens against 184,772 for Sonnet, with the same median of 7 turns, and that its median cache write count was 45,364 against 25,579. Reveneau's reading is limited to what the counts say: the unchanged start of each request, which the Opus runs re-read every time, was larger. Finding the cause would need a comparison of the transcripts, which was not done.
What happened in the 159-second Opus run?
One of the five Opus runs took 159 seconds of wall-clock time and 8 turns, against 18 to 22 seconds and 7 turns for the other four, and cost $0.612 against $0.488 to $0.490 for the others. The cause was not recorded. The run passed. The median of $0.489 reports the usual run, and the table gives the minimum and maximum so the unusual run is visible. Across all 45 benchmark runs, this was the slowest and the most costly.
Did the benchmark measure Opus 5.5, Sonnet 5.5 or Fable 5.1?
No. The installed Claude Code was 2.1.118 and its aliases resolved to claude-sonnet-4-6, claude-opus-4-7 and claude-haiku-4-5-20251001. Anthropic's current documentation, read on 4 October 2026, names Opus 5.5, Sonnet 5.5 and Fable 5.1 and gives them different prices and defaults. Prices and behaviour for those models were not measured in this run, and the ratios on this page apply only to the three models that were run.
Which list prices apply to the three models in the benchmark?
Anthropic's pricing page, read on 4 October 2026, lists per million tokens: Claude Haiku 4.5 at $1 input, $2 one-hour cache write, $0.10 cache read and $5 output; Claude Sonnet 4.6 at $3, $6, $0.30 and $15; Claude Opus 4.7 at $5, $10, $0.50 and $25. These are Anthropic's list prices. Claude Code computes its total_cost_usd figure from a bundled price table at list price, so these are the rates behind the benchmark's dollar figures.
Did the benchmark change the effort level between models?
No. No --effort flag was passed in the model experiment, so each model ran at whatever default effort level Claude Code 2.1.118 applied to it. Anthropic's current model configuration page, read on 4 October 2026, states per-model defaults, and Reveneau did not verify that those defaults applied in version 2.1.118, so this page does not state which level each model ran at. The effort experiment, on Sonnet only, set the level with the flag.
Does the model experiment show which model to use for coding?
No. It shows what three models cost on one invented task of one failing test, five runs each, on 4 October 2026 with Claude Code 2.1.118. Anthropic's costs page, read the same day, says Sonnet handles most coding tasks well and costs less than Opus, and to reserve Opus for complex architectural decisions or multi-step reasoning. Reveneau's measurement agrees with the direction of that advice on this task, and a harder task could give a different result.
Why does the model table repeat the Sonnet row from the CLAUDE.md experiment?
Because the same five runs serve both experiments. The no-CLAUDE.md Sonnet setting at default effort is setting A of the CLAUDE.md experiment, the Sonnet row of the model experiment and the default row of the effort experiment. Running it once and reusing it kept the benchmark at nine settings and 45 runs. The row reads the same everywhere: $0.161 median, $0.161 to $0.163 range, 184,772 cache read tokens, 7 turns.
References
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Model configuration (code.claude.com), read 4 October 2026
- Anthropic, Pricing (platform.claude.com), read 4 October 2026
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Track cost and usage (code.claude.com), read 4 October 2026
- Anthropic, CLI reference (code.claude.com), read 4 October 2026
More in Experiments
The CLAUDE.md length experiment: 301 lines against none
In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times with no CLAUDE.md and five times with a 301-line CLAUDE.md of project rules unrelated to the task (62,746 characters; 342 lines when the blank lines between sections are counted), on claude-sonnet-4-6 at default effort with Claude Code 2.1.118. The median cost rose from $0.161 to $0.239 by Claude Code's own list-price figure on a subscription account, which is $0.078 or 48 percent more by Reveneau's arithmetic. Cache write tokens rose by 14,487 and cache read tokens by 71,678. All ten runs passed. One task is one task, the figures are medians of five, and the experiment measured one file against none, with nothing in between.
The effort level experiment: low, medium and high on Sonnet 4.6
In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times at each of three effort levels, passed with the --effort flag, on claude-sonnet-4-6 with no CLAUDE.md and Claude Code 2.1.118 on a Claude subscription. The median costs by Claude Code's own list-price figure were $0.150 at low, $0.162 at medium and $0.170 at high. By Reveneau's arithmetic the gap between low and high is $0.020, or 13 percent of the low median, and the ranges overlap: one low run cost $0.208, above every high run. The difference in medians came from turns, 6 against 7; output token medians are within 33 tokens of each other. All 15 runs passed. One task is one task, and the figures are medians of five.
The output-trimming hook experiment: a test that prints 3,000 lines
In Reveneau's Claude Code token benchmark of 4 October 2026, a second test printing 3,000 lines was added and the task ran five times with no hook and five times through Anthropic's example PreToolUse hook that filters test output, on claude-sonnet-4-6 with Claude Code 2.1.118 on a Claude subscription. The hook runs cost more: a median of $0.201 against $0.159 by Claude Code's own list-price figure, with a median of 10 turns against 7. The transcripts show why: in every direct run the model itself added | tail -20 to the test command, keeping the last 20 lines, and in the hook runs the filter left no output on a pass, so the model ran the tests again. One task is one task; the figures are medians of five.