Limits

What this Claude Code token benchmark cannot show

Reveneau's Claude Code token benchmark of 4 October 2026 measured one invented task of one failing test in a 30-line project, five runs per setting, with Claude Code 2.1.118 on a Claude subscription, so it cannot show what a long session costs, what /clear or /compact saves, what happens when the cache expires, how much of a plan's allowance a run uses, or what newer models cost. Its dollar figures are Claude Code's own list-price estimates, its medians are medians of five, and its CLAUDE.md experiment measured a file the task did not need. This page lists each limit, says which claims the run supports and which it does not, and names the experiments planned for a later dated run.

Published October 4, 2026. Editorial.

Key takeaways

  • One invented task of one failing test in a 30-line project is one task; nothing in the benchmark is a statistic about Claude Code in general, and every median is a median of five runs.
  • The task took a median of 6 to 10 turns per setting, so the run cannot show what a long session costs, what /clear or /compact saves, or what happens when the prompt cache expires; Anthropic's costs page lists long context and cache misses among the reasons a session open for hours uses more of a plan.
  • The dollar figures are Claude Code's local list-price estimates on a subscription account, and no Claude Code output reports how much of a plan's five-hour or weekly allowance a run used.
  • The installed Claude Code was 2.1.118 and its aliases resolved to Sonnet 4.6, Opus 4.7 and Haiku 4.5; Anthropic's current documentation gives Opus 5.5, Sonnet 5.5 and Fable 5.1 different prices, defaults and cache behaviour, none of which was measured.
  • Four planned experiments were not run: tool servers, one long session against /clear between tasks, a subagent against the main session, and inside against after the cache lifetime; they are planned for a later dated run with no schedule.

Every experiment page in Reveneau's Claude Code token benchmark ends with what it cannot show. This page collects those limits in one place, adds the ones that apply to the whole run, and gives a table of claims a reader might make with a yes or no beside each. Reveneau is an AI software development consultancy whose code is written by AI; it measures token use because tokens are a running cost of every build, and it publishes the limits beside the figures because a figure without its limits becomes a different claim. Reveneau is independent of Anthropic. The method is on the method page.

The words used on this page

A token is a piece of text the model reads or writes. The context window is all the text the model reads in one request. Claude Code sends the whole conversation again on every request; the service stores the unchanged start, and tokens stored this way are cache write tokens, while tokens read back from the store on a later request are cache read tokens [2]. CLAUDE.md is a file of instructions loaded at the start of every session. The effort level is a setting for how much reasoning the model does before each answer. A hook is a shell command Claude Code runs on its own at a fixed point. A non-interactive run is claude -p with a prompt: Claude Code does the task and exits. A turn is one request to the model and its reply. A median is the middle value of five sorted runs.

One task is one task

The task was invented for the run: one failing test in a 30-line Node.js project, fixed by one line. Five runs per setting gave 45 recorded runs across nine settings, all on 4 October 2026 with Claude Code 2.1.118 on a Claude subscription, with aliases resolved to claude-sonnet-4-6, claude-opus-4-7 and claude-haiku-4-5-20251001. Nothing here is a statistic about Claude Code in general. The medians are medians of five, and each table gives the range so a reader can see how far the five runs spread. A different task would give different medians, and the pages say so wherever a figure appears.

The runs were short

The task took a median of 6 to 10 turns per setting, and no single run took more than 15. Every run finished in under 160 seconds, and the median wall-clock time (the time from start to finish on an ordinary clock) in every setting was 21 seconds or less. That rules out three things the run cannot measure.

It cannot show what a long session costs. Anthropic's costs page, read on 4 October 2026, says a session open for hours can use more of a plan than the activity suggests, and lists first among the reasons that Claude Code sends the full conversation with every request and re-reads the history at the cached rate [1]. A 7-turn run has no such history.

It cannot show what /clear or /compact saves. Each run was a new session that ended with the task, so there was never a second task to clear before and never a conversation long enough to compact. Anthropic's documentation says /clear costs nothing and that /compact sends a separate request that reads the conversation it summarises, at the cached rate while the cache is warm and in full after a break [1][2]. None of that happened in the run. The sister page on /clear, /compact and /rewind covers Anthropic's guidance.

It cannot show what happens when the cache expires. Anthropic's prompt caching page says the main conversation's cache lifetime is one hour on a subscription within plan usage and five minutes on an API key, and that the first request after a longer gap reprocesses the full context [2]. Every run finished in under 160 seconds, inside the one-hour lifetime, so no run paid that cost. The experiment that was planned for this, inside against after the cache lifetime, was not run. The prompt caching guide covers the two lifetimes.

Anthropic's costs page lists long context and cache misses first among the reasons a session open for hours uses more of a plan than the activity suggests [1]. The benchmark measured neither. Its figures describe the cost of a short task, which is a different question.

The dollar figures are estimates on a subscription

Every dollar figure is Claude Code's total_cost_usd, computed locally from a bundled price table at list price. Anthropic's costs page, read on 4 October 2026, says the Session block in /usage shows API token usage intended for API users, and that Pro and Max subscribers have usage included in their subscription [1]. The benchmark account is a subscription within its plan usage, so there is no invoice for these runs; the figures are the comparable quantity across settings and nothing more.

What the figures do not show is how much of a plan's five-hour or weekly allowance a run used. No Claude Code output reports that, so the run cannot say what fraction of a Pro or Max plan a $0.161 task represents. The usage limits guide covers what counts against an allowance, and the team costs guide covers how a subscription and an API key differ in what a dollar figure means.

One more consequence of the subscription: every cache write was a one-hour write, as the method page explains. A reader on an API key would see five-minute writes by default, priced differently by Anthropic's pricing page [3], so cache write figures from the two setups are not directly comparable.

The version and the models

The installed Claude Code was 2.1.118, and its aliases resolved to Sonnet 4.6, Opus 4.7 and Haiku 4.5. Anthropic's current documentation, read on 4 October 2026, names Opus 5.5, Sonnet 5.5 and Fable 5.1 and gives them different list prices [3], a default effort of medium on Opus 5.5 and Sonnet 5.5 against high on Sonnet 4.6 and xhigh on Opus 4.7 [4], and a cache that survives an effort change on those models with an API key or a subscription, while on most models an effort change recomputes the entire request [2]. Each of those would change at least one experiment, and none was measured.

Token counts do not transfer either. Anthropic's pricing page says Claude 4.7 and later models use a newer tokenizer (the rule a model uses to split text into tokens) that produces more tokens for the same text, which Anthropic puts at 30 percent more with the exact increase depending on the content and workload, and that Sonnet 4.6 and earlier use the previous tokenizer [3]. The benchmark's Sonnet counts cannot be read as what the same text would count on a later model. The model choice page gives the prices that applied to the three models that were run.

Anthropic's costs page also says Claude Code regularly receives updates that may change how features work, including cost reporting [1]. A later version may count or price differently, which is why the version appears on every page that gives a number.

The CLAUDE.md the task did not need

The CLAUDE.md experiment measured one 301-line file of rules unrelated to the task against no file. It measured the cost of text that is loaded on every request and never used, and it found that cost on this task: $0.078 more at the median, 48 percent of the no-CLAUDE.md median by Reveneau's arithmetic. It did not measure a CLAUDE.md the task needs. Such a file could change the number of turns, by telling the model where to look or by adding steps, and on this task turns drive the cost, so the run cannot say what a useful file costs. It also tested no length between zero and 301 lines, so it cannot say how the cost changes between zero and 301 lines, and it cannot test Anthropic's 200-line guidance.

The experiments not run

The plan had more experiments than the run. Four were not run: experiment 2, tool servers, which would compare a session with MCP servers (programs that give Claude Code extra tools) connected against one without; experiment 3, one long session against /clear between tasks; experiment 7, a subagent (a separate conversation the main one starts for part of the work) against the main session for the same work; and experiment 8, a run inside the cache lifetime against one after it. Each needs a longer or repeated session than the one-task method gives, which is why they were set aside. They are planned for a later dated run. No schedule is given, because a date would be a promise, and the first run's data will stay published beside whatever follows.

Three further things were not measured and are worth naming. The hook experiment ran one filter and did not run a version that keeps the summary line. The effort experiment did not run max, and ran no level on Haiku or Opus. No experiment had a run that failed, so nothing here measures what a setting does to quality; the pages compare the cost of settings that all did the job.

Why the limits are published beside the figures

A number is quoted more often than the sentence around it. The 48 percent from the CLAUDE.md experiment is true of one task with a median of 7 or 8 turns on one day with one version and one model, and false as a general statement about CLAUDE.md files. If the page gave only the number, a reader would quote the general statement, and Reveneau would have published something it did not measure. Every experiment page therefore repeats that one task is one task and that the figures are medians of five, names the version and the models, and says where Reveneau did not verify a cause and offers no guess. That is also why every raw row is published: a reader who disagrees with a median can compute their own from runs.json.

How to quote the numbers

A claim a reader might make Does this run support it
On one task of one failing test, run five times on 4 October 2026 with Claude Code 2.1.118 and claude-sonnet-4-6, a 301-line unused CLAUDE.md raised the median cost from $0.161 to $0.239 by Claude Code's list-price figure Yes, with every part of that sentence kept
A long CLAUDE.md costs 48 percent more No
On that task, Haiku 4.5's median cost was 37 percent of Sonnet 4.6's and Opus 4.7's was 3.0 times Sonnet 4.6's, medians of five, by Reveneau's arithmetic Yes
Haiku costs less than Sonnet for coding No; one task, three resolved models, one day
Opus reads more cache tokens than Sonnet No; measured once on one task, cause not verified
Low effort saves 13 percent No; one task, medians of five, and the ranges overlap
Anthropic's test-output hook makes runs cost more No; one filter, one task, and the model piped output through tail on its own
On that task, the filtered runs took a median of 10 turns against 7 and cost $0.201 against $0.159 by Claude Code's figure Yes
Claude Code costs $0.16 per task No
A Claude Code run uses a certain share of a Pro or Max allowance No; nothing in the run reports that
The result applies to Opus 5.5, Sonnet 5.5 or Fable 5.1 No; none was run
All 45 recorded runs passed, and the run cost $9.15 in total by Claude Code's list-price figure Yes
Reveneau's own projects cost a stated amount in tokens No; nothing on these pages says anything about that

The rule in the table is the rule for the whole guide: a figure keeps its task, date, version, model and run count in the same sentence, or it is a different claim. A reader who wants a figure for their own task can get one with the do-it-yourself page, and Anthropic's own published cost figures, which the team costs guide covers, are Anthropic's statements about its own customers and a different type of number again.

Common questions

Why does Reveneau publish the limits of its benchmark beside the results?

Because a figure quoted without its limits becomes a different claim. The benchmark's 48 percent for a 301-line CLAUDE.md is true of one task with a median of 7 or 8 turns on 4 October 2026 with Claude Code 2.1.118 and claude-sonnet-4-6, and false as a general statement about CLAUDE.md files. Reveneau's position is that a measurement is only useful when a reader can see what it covers, so every experiment page states its own limits and this page collects them in one place.

Can the benchmark figures be quoted as what Claude Code costs?

No. They are what one invented task of one failing test in a 30-line project cost, five runs per setting, on one day, one version and three resolved models, by Claude Code's own list-price figure on a subscription account. The median Sonnet run cost $0.161 on that task; a task twice as long, a different model, a CLAUDE.md the task needs, or a different version would each give a different figure. Quote a figure with its task, date, version and model, or do not quote it.

Why can the benchmark say nothing about long sessions?

Because every setting had a median of 6 to 10 turns, no run took more than 15, and every run took under 160 seconds. Anthropic's costs page, read on 4 October 2026, lists long context, where the whole conversation is re-sent on every request, and cache misses after a break longer than the cache lifetime, among the reasons a session open for hours uses more of a plan. Neither happens in a 7-turn run, so the benchmark measured none of what makes a long session cost what it does.

Does the benchmark show what /clear or /compact saves?

No. Each run was a new session that ended when the task did, so there was never a second task to clear before, never a conversation long enough to compact, and never a summary request to cost. Anthropic's documentation, read on 4 October 2026, says /clear costs nothing and that /compact sends a separate request that reads the conversation it summarises. The planned experiment comparing one long session against /clear between tasks was not run and is planned for a later dated run.

Do the benchmark's dollar figures show how much of a subscription allowance a run used?

No. The figures are Claude Code's local estimates at list price, which Anthropic's costs page, read on 4 October 2026, says are intended for API users, while subscribers have usage included in their plan. No Claude Code output reports what share of a plan's five-hour or weekly allowance a run consumed, so the benchmark cannot translate $0.161 into a fraction of a Pro or Max plan. The usage limits guide covers what counts against an allowance.

Do the benchmark results apply to Opus 5.5, Sonnet 5.5 or Fable 5.1?

No. Claude Code 2.1.118 resolved its aliases to claude-sonnet-4-6, claude-opus-4-7 and claude-haiku-4-5-20251001, and only those three were run. Anthropic's documentation, read on 4 October 2026, gives the newer models different list prices, a default effort of medium on Opus 5.5 and Sonnet 5.5 against high on Sonnet 4.6, and a cache that survives an effort change on those models. Each of those would change at least one experiment's result, and none was measured.

Can the benchmark's token counts be compared with a run on Opus 5.5 or Sonnet 5.5?

No. A token count is specific to the way a model splits text. Anthropic's pricing page, read on 4 October 2026, says Claude 4.7 and later models use a newer tokenizer that produces more tokens for the same text, which Anthropic puts at 30 percent more with the exact increase depending on the content, and that Sonnet 4.6 and earlier use the previous one. The benchmark's Sonnet token counts therefore cannot be transferred to a 4.7 or later model as if the text would count the same.

What would a CLAUDE.md the task needs change in the benchmark?

The number of turns, in either direction, and with it the cost. The CLAUDE.md experiment used 301 lines of rules unrelated to the task, so the file added tokens to every request and changed nothing about how the model worked. A file with a rule the task needs could make the model take fewer turns, by telling it where to look, or more, by adding steps. On this task turns drive the cost, so the run cannot say what a useful file costs.

Which planned experiments did the benchmark not run?

Four from the plan: experiment 2 on tool servers, experiment 3 comparing one long session against /clear between tasks, experiment 7 comparing a subagent against the main session, and experiment 8 comparing a run inside the cache lifetime against one after it. Each needs a longer or repeated session than the one-task method provides. They are planned for a later dated run, and the first run's data at /benchmarks/claude-code-tokens/2026-10-04/ will stay published beside it.

When will Reveneau run the next benchmark?

A later dated run is planned, with no schedule. Reveneau does not state a date because a date would be a promise it has not measured its ability to keep. The next run will have its own date in its path, as the first does at /benchmarks/claude-code-tokens/2026-10-04/, and its pages will state the Claude Code version and the resolved model names, so two runs can be compared without assuming the product stayed the same between them.

How should I cite a figure from the Reveneau benchmark?

Name the task, the date, the version, the model and the run count in the same sentence as the figure, and link the raw data. For example: on one task of one failing test in a 30-line project, run five times per setting on 4 October 2026 with Claude Code 2.1.118 and claude-sonnet-4-6, a 301-line unused CLAUDE.md raised the median cost from $0.161 to $0.239 by Claude Code's own list-price figure, which is Reveneau's arithmetic of 48 percent. Leave out any of those and the claim changes.

Can the benchmark show whether a setting changes the quality of Claude Code's work?

No. All 45 recorded runs passed, and the task has one correct outcome, so every table compares the cost of settings that all did the job. A setting that lowered cost and raised the failure rate would show as a lower cost and a lower pass count, and no setting did that here. A task with several acceptable outcomes, or one hard enough that some runs fail, would be needed to measure quality, and the benchmark did not include one.