Experiments

The effort level experiment: low, medium and high on Sonnet 4.6

In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times at each of three effort levels, passed with the --effort flag, on claude-sonnet-4-6 with no CLAUDE.md and Claude Code 2.1.118 on a Claude subscription. The median costs by Claude Code's own list-price figure were $0.150 at low, $0.162 at medium and $0.170 at high. By Reveneau's arithmetic the gap between low and high is $0.020, or 13 percent of the low median, and the ranges overlap: one low run cost $0.208, above every high run. The difference in medians came from turns, 6 against 7; output token medians are within 33 tokens of each other. All 15 runs passed. One task is one task, and the figures are medians of five.

Published October 4, 2026. Editorial.

Key takeaways

  • On the benchmark task the median cost was $0.150 at low effort, $0.162 at medium and $0.170 at high, by Claude Code's own list-price figure on a subscription account, medians of five runs on claude-sonnet-4-6.
  • By Reveneau's arithmetic the gap between low and high is $0.020, which is 13 percent of the low median, and one low run cost $0.208, above every high run, so the ranges overlap.
  • The median output token counts were 644, 677 and 655, within 33 tokens of each other; the difference in medians came from turns, 6 at low against 7 at medium and high.
  • The task is too small to need extended reasoning, so the experiment cannot show what effort does on a task that does.
  • Anthropic's documentation, read on 4 October 2026, says thinking tokens are billed as output tokens and that Sonnet 4.6 offers the levels low, medium, high and max.

This is the third experiment in Reveneau's Claude Code token benchmark. It asks what the effort level costs on a task that does not need much reasoning. Reveneau is an AI software development consultancy whose code is written by AI, so a setting that changes the cost of every request is worth measuring before it is set as a default across builds. Reveneau is independent of Anthropic. The full method is on the method page.

The words used on this page

A token is a piece of text the model reads or writes. The context window is all the text the model reads in one request. Claude Code sends the whole conversation again on every request; the service stores the unchanged start, and tokens stored this way are cache write tokens, while tokens read back from the store on a later request are cache read tokens [3]. CLAUDE.md is a file of instructions loaded at the start of every session; this experiment used none. The effort level is a setting that tells the model how much reasoning to do before each answer: Anthropic's model configuration page, read on 4 October 2026, says effort levels control adaptive reasoning, which lets the model decide whether and how much to think on each step, and that lower effort is faster and costs less for straightforward tasks while higher effort gives deeper reasoning for complex problems [1]. A hook is a shell command Claude Code runs on its own at a fixed point; none was used. A non-interactive run is claude -p with a prompt: Claude Code does the task and exits. A turn is one request to the model and its reply. A median is the middle value of five sorted runs.

What was compared

Three settings, each run five times on claude-sonnet-4-6: --effort low, --effort medium and --effort high. Anthropic's CLI reference, read on 4 October 2026, says the --effort flag sets the effort level for the current session, overrides the saved settings for that session and does not persist [4]. The model configuration page lists the levels each model offers: low, medium, high and max for Opus 4.6 and Sonnet 4.6, and xhigh in addition on Opus 4.7, Opus 4.8, Opus 5, Sonnet 5, Opus 5.5 and Sonnet 5.5 [1]. The benchmark did not run max.

The fourth row of the table is the no-CLAUDE.md Sonnet setting from the CLAUDE.md experiment, which passed no --effort flag. It is reported separately, and the reason is explained below.

Everything else was constant: the task of one failing test in a 30-line Node.js project, the same prompt, no CLAUDE.md, the flags on the method page, Claude Code 2.1.118 on a Claude subscription, and a fresh copy of the folder for every run. All runs were on 4 October 2026.

The result

Every figure is Reveneau's own measurement from that run, a median of five, with the dollar figure being Claude Code's own list-price figure on a subscription account. The rows are in runs.json.

Effort Pass Cost median (min to max) Cache read median Cache write median Output median Turns median Seconds median
low 5 of 5 $0.150 ($0.149 to $0.208) 150,161 25,515 644 6 16
medium 5 of 5 $0.162 ($0.149 to $0.201) 184,771 25,600 677 7 16
high 5 of 5 $0.170 ($0.161 to $0.174) 187,716 25,636 655 7 17
default (no flag), from experiment 1 5 of 5 $0.161 ($0.161 to $0.163) 184,772 25,579 659 7 21

All 15 flagged runs passed, and so did the five default runs.

Reveneau's reading

The effort level moved the median cost by $0.020 between low and high, from $0.150 to $0.170, which is 13 percent of the low median by Reveneau's arithmetic. The ranges overlap: one low run cost $0.208, which is above every high run, and the lowest medium run and the lowest low run both cost $0.149. Five runs can order the medians on this task; they cannot separate the settings run by run.

The difference in medians came from turns. Low effort had a median of 6 turns and medium and high had 7. Each turn re-sends the whole conversation and reads the unchanged start from the cache, so one more turn is one more read of the growing unchanged start; the median cache read count was 150,161 at low, 184,771 at medium and 187,716 at high. The cache write medians were within 121 tokens of each other, 25,515 to 25,636, by Reveneau's arithmetic, which fits a run that builds the same unchanged start at every level.

The output token medians were 644, 677 and 655, within 33 tokens of each other. Anthropic's costs page, read on 4 October 2026, says thinking tokens are billed as output tokens and that the default thinking budget can be tens of thousands of tokens per request depending on the model [2]. If the model had reasoned at length at high effort, that would have shown in the output column. It did not, because the task is one line and a test. The experiment therefore measures what the flag costs when the task does not use the reasoning it allows: on this task, one turn's worth of cache reads.

Why the default row is separate

Anthropic's model configuration page, read on 4 October 2026, says Claude Code resolves the effort level in order: an explicit choice such as --effort, then the saved settings, then the model's default, which is high on every model that supports effort except Opus 5.5 and Sonnet 5.5 at medium and Opus 4.7 at xhigh [1]. That page describes the current version. Reveneau did not verify which level Claude Code 2.1.118 applied to claude-sonnet-4-6 when no flag was passed, and the default row's median of $0.161 is reported beside the three flagged rows without labelling it as any of them. The default row also differs from the flagged rows in its median wall-clock time (the time from start to finish on an ordinary clock), 21 seconds against 16 or 17, and the benchmark did not investigate that.

What the task cannot show about effort

The task is too small to need extended reasoning. The fix is one line, the test confirms it, and the median output was 644 to 677 tokens at every level. Anthropic's model configuration page describes high as the level for work where verification matters or edge cases are likely, and max as the level for hard problems the model works through on its own, with a warning that it may show diminishing returns [1]. None of that applies to this task, so the experiment cannot show what effort does on a task that does use it. On a task where the model reasons at length, the output column would move, and the cost with it.

The experiment also did not run max, and did not run any level on Haiku or Opus. The model choice experiment ran those two models at their defaults only.

Changing effort mid-task was not measured

Every run set its level at launch and kept it. Anthropic's prompt caching page, read on 4 October 2026, says that on most models each effort level has its own cache, so changing effort mid-session recomputes the entire request, and that on Opus 5.5, Sonnet 5.5 and Fable 5.1 with an API key or a Claude subscription the cache stays intact [3]. No run in the benchmark paid or avoided that cost; the prompt caching guide covers it.

What Reveneau recommends

Reveneau's recommendation, drawn from Anthropic's documentation and consistent with this measurement: pick the effort level for the task at the start of the session, and use a lower level for a small, well-defined change where you review each result. Anthropic's costs page says that for simpler tasks where deep reasoning is not needed, lowering the effort level reduces cost [2], and on this task low effort passed every run in a median of 6 turns for $0.150 by Claude Code's figure. The saving on this task was $0.020 per run between low and high, and single runs varied by more than that, so on a task this small the effort level is a smaller cost decision than the model, where the model choice experiment measured a difference of 3.0 times between Sonnet and Opus. The sister page on which model and effort level to choose covers Anthropic's guidance for each level, and the token usage guide explains why the turn count set the cost on this task.

What this experiment cannot show

One invented task of one failing test in a 30-line project is one task. The medians are medians of five, and nothing here is a statistic about Claude Code in general. The model was claude-sonnet-4-6 as resolved by Claude Code 2.1.118 on 4 October 2026; the same version resolved opus to claude-opus-4-7 and haiku to claude-haiku-4-5-20251001, and no effort level was run on either. Anthropic's current documentation names Sonnet 5.5 with a different default effort and a different cache behaviour on an effort change, and that model was not measured. The dollar figures are Claude Code's own list-price estimates on a subscription account, and they do not show how much of a plan's allowance a run used, which no Claude Code output reports. The limits page lists everything the run cannot show, and a later dated run is planned, with no schedule.

Common questions

Which effort levels did the Reveneau benchmark run on Sonnet 4.6?

The benchmark ran low, medium and high, each passed with the --effort flag at launch, five runs each on claude-sonnet-4-6 with no CLAUDE.md and Claude Code 2.1.118 on 4 October 2026. The default setting with no flag, from the CLAUDE.md experiment, is reported as a fourth row. Anthropic's model configuration page, read the same day, lists low, medium, high and max for Opus 4.6 and Sonnet 4.6; max was not run.

How much did the effort level change the median cost on the benchmark task?

The median cost was $0.150 at low, $0.162 at medium and $0.170 at high, by Claude Code's own list-price figure on a subscription account. By Reveneau's arithmetic the difference between low and high is $0.020, which is 13 percent of the low median. These are medians of five runs of one task of one failing test in a 30-line project, so the figure describes that task and nothing wider.

Why do the effort experiment's cost ranges overlap?

The five low-effort runs ranged from $0.149 to $0.208, the medium runs from $0.149 to $0.201 and the high runs from $0.161 to $0.174. One low run at $0.208 cost more than every high run. The medians differ by $0.020 while single runs of the same setting differ by up to $0.059 by Reveneau's arithmetic on the low range, so five runs can order the medians and cannot separate the settings run by run.

Where did the cost difference between low and high effort come from on the benchmark task?

From the turn count. Low effort had a median of 6 turns and medium and high had 7. Each turn re-sends the whole conversation and reads the unchanged start from the cache, so one more turn adds one more read of the growing conversation: the median cache read count was 150,161 at low against 187,716 at high. Output token medians were 644, 677 and 655, within 33 tokens of each other, so the extra reasoning that effort controls did not show up as output on this task.

Did a higher effort level produce more output tokens on the benchmark task?

No measurable amount. The median output counts were 644 tokens at low, 677 at medium and 655 at high, within 33 tokens of each other by Reveneau's arithmetic. Anthropic's costs page, read on 4 October 2026, says thinking tokens are billed as output tokens, so more reasoning would appear in this column. On a task of one failing test the model did not need to reason at length at any level, which is what the near-equal output counts record.

Why is the default effort row listed separately in the effort experiment?

Because Reveneau did not verify which level Claude Code 2.1.118 applied when no flag was passed. Anthropic's model configuration page, read on 4 October 2026, says the default is high on every model that supports effort, except Opus 5.5 and Sonnet 5.5 at medium and Opus 4.7 at xhigh, but that page describes the current version. The default row, from the CLAUDE.md experiment, had a median of $0.161, and the page reports it beside the flagged rows without labelling it as any of them.

Did the benchmark test the max effort level?

No. Anthropic's model configuration page, read on 4 October 2026, lists low, medium, high and max for Sonnet 4.6 and describes max as the deepest reasoning level, for hard problems, with a warning that it may show diminishing returns. The benchmark ran low, medium and high. A task of one failing test gave the three levels it ran output medians within 33 tokens of each other, so there was no reason to expect max to show anything on it; a later dated run on a harder task could include it.

Can the effort experiment show what effort does on a hard task?

No. The task is too small to need extended reasoning: the median output was 644 to 677 tokens at every level and the fix is one line. Anthropic's documentation says higher effort provides deeper reasoning for complex problems, and this task is not one. The experiment measures what the flag costs when the task does not use the reasoning it allows, which is one turn's difference on this task. What effort does on a task that needs it was not measured.

Does the benchmark report thinking tokens separately from output tokens?

No. The benchmark read the output_tokens count from each json result, and Anthropic's costs page, read on 4 October 2026, says thinking tokens are billed as output tokens, so any reasoning the model did is inside that count. The output medians at the three levels were 644, 677 and 655. The json result does not break output into thinking and answer, and the benchmark did not try to.

Did changing the effort level affect the prompt cache in the benchmark?

No, because no run changed it. Every run set its effort level with the --effort flag at launch and kept it, so each run built and read its own cache at one level. Anthropic's prompt caching page, read on 4 October 2026, says that on most models each effort level has its own cache, so changing effort mid-session recomputes the entire request, and that Opus 5.5, Sonnet 5.5 and Fable 5.1 keep the cache; neither case arose in these runs.

Did every run in the effort experiment pass?

Yes. All five runs at low, all five at medium and all five at high passed, graded by re-running node --test test/ledger.test.js outside Claude Code and treating exit code 0 as a pass. The lowest effort level did the task every time in a median of 6 turns. The experiment therefore compares the cost of settings that all did the job and says nothing about whether effort changes the quality of a result on a task with more than one acceptable outcome.

More in Experiments

The CLAUDE.md length experiment: 301 lines against none

In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times with no CLAUDE.md and five times with a 301-line CLAUDE.md of project rules unrelated to the task (62,746 characters; 342 lines when the blank lines between sections are counted), on claude-sonnet-4-6 at default effort with Claude Code 2.1.118. The median cost rose from $0.161 to $0.239 by Claude Code's own list-price figure on a subscription account, which is $0.078 or 48 percent more by Reveneau's arithmetic. Cache write tokens rose by 14,487 and cache read tokens by 71,678. All ten runs passed. One task is one task, the figures are medians of five, and the experiment measured one file against none, with nothing in between.

The model choice experiment: Haiku, Sonnet and Opus on one task

In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times on each of three models, with no CLAUDE.md and default effort, using Claude Code 2.1.118 on a Claude subscription. The aliases resolved to claude-haiku-4-5-20251001, claude-sonnet-4-6 and claude-opus-4-7. The median costs by Claude Code's own list-price figure were $0.060, $0.161 and $0.489. By Reveneau's arithmetic, Haiku's median was 37 percent of Sonnet's and Opus's was 3.0 times Sonnet's. All 15 runs passed. Haiku took more turns, a median of 9 against 7. One task is one task, the figures are medians of five, and newer models were not measured.

The output-trimming hook experiment: a test that prints 3,000 lines

In Reveneau's Claude Code token benchmark of 4 October 2026, a second test printing 3,000 lines was added and the task ran five times with no hook and five times through Anthropic's example PreToolUse hook that filters test output, on claude-sonnet-4-6 with Claude Code 2.1.118 on a Claude subscription. The hook runs cost more: a median of $0.201 against $0.159 by Claude Code's own list-price figure, with a median of 10 turns against 7. The transcripts show why: in every direct run the model itself added | tail -20 to the test command, keeping the last 20 lines, and in the hook runs the filter left no output on a pass, so the model ran the tests again. One task is one task; the figures are medians of five.