The CLAUDE.md length experiment: 301 lines against none
In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times with no CLAUDE.md and five times with a 301-line CLAUDE.md of project rules unrelated to the task (62,746 characters; 342 lines when the blank lines between sections are counted), on claude-sonnet-4-6 at default effort with Claude Code 2.1.118. The median cost rose from $0.161 to $0.239 by Claude Code's own list-price figure on a subscription account, which is $0.078 or 48 percent more by Reveneau's arithmetic. Cache write tokens rose by 14,487 and cache read tokens by 71,678. All ten runs passed. One task is one task, the figures are medians of five, and the experiment measured one file against none, with nothing in between.
Published October 4, 2026. Editorial.
Key takeaways
- With the 301-line CLAUDE.md the median cost was $0.239 against $0.161 without it, a rise of $0.078 or 48 percent of the no-CLAUDE.md median, by Reveneau's arithmetic on medians of five runs.
- Cache write tokens rose by 14,487 per run and cache read tokens by 71,678, which fits the file being written to the cache once and read back on every later request, as Anthropic's prompt caching page describes.
- The task did not need anything in the file: it held 20 sections of 14 numbered rules each, all unrelated to the failing test, and all ten runs passed.
- Anthropic's documentation, read on 4 October 2026, says to keep CLAUDE.md under 200 lines; the experiment measured 301 lines against zero and tested no length in between.
- The run cannot say what a CLAUDE.md the task does need would cost, because such a file could change the number of turns, and turns drive the cost on this task.
This is the first of four experiments in Reveneau's Claude Code token benchmark. The question it asked is narrow: what does a long CLAUDE.md cost on a task that does not use it? Reveneau is an AI software development consultancy whose code is written by AI, so token use is a running cost of every build, and a file that is loaded on every request is a cost that is easy to stop noticing. Reveneau is independent of Anthropic, and the method is on the method page.
The words used on this page
A token is a piece of text that the model reads or writes; Anthropic's pricing page, read on 4 October 2026, estimates one token at 4 characters of English [4]. The context window is all the text the model reads in one request. Claude Code sends the whole conversation again on every request; the service stores the unchanged start of it, and tokens stored this way are cache write tokens, while tokens read back from that store on a later request are cache read tokens [3]. CLAUDE.md is a file of instructions that Claude Code reads at the start of every session and includes in every request [2]. The effort level is a setting for how much reasoning the model does before each answer; this experiment left it at the default. A hook is a shell command Claude Code runs on its own at a fixed point; none was used here. A non-interactive run is claude -p with a prompt: Claude Code does the task and exits. A turn is one request to the model and its reply, so a run with several tool calls has several turns. A median is the middle value of five sorted runs.
What was compared
Setting A had no CLAUDE.md in the task folder. Setting B had a CLAUDE.md of 301 lines of text and 62,746 characters, arranged as a title and 20 headed sections of 14 numbered project rules each, none of which concerned the task; counting the blank lines between sections, the file is 342 lines long. The file is published as task/CLAUDE-long.md, and the script run.sh copies it into the fresh task folder for setting B only.
Everything else was the same in both settings: the task of one failing test in a 30-line Node.js project, the same prompt, claude-sonnet-4-6 at the default effort level, Claude Code 2.1.118 signed in with a Claude subscription, the flags listed on the method page, and a fresh copy of the folder for each run. Five runs per setting, on 4 October 2026.
The content of the file was unrelated to the task on purpose. The question was the cost of instructions that are loaded on every request and never used. Anthropic's costs page, read on 4 October 2026, describes exactly that situation: a CLAUDE.md with detailed instructions for specific workflows keeps those tokens present even during unrelated work, which is why it suggests moving such instructions into skills that load only when invoked [1]. A file of rules the task did need would have mixed two effects, the cost of the text and any change in how the model worked, and five runs could not have separated them.
The result
Every figure is Reveneau's own measurement from the run of 4 October 2026, read from the json result of each run, and every row is a median of five. The dollar figure is Claude Code's own list-price figure on a subscription account. The raw rows are in runs.json.
| Setting | Pass | Cost median (min to max) | Cache read median | Cache write median | Output median | Turns median | Seconds median |
|---|---|---|---|---|---|---|---|
| No CLAUDE.md | 5 of 5 | $0.161 ($0.161 to $0.163) | 184,772 | 25,579 | 659 | 7 | 21 |
| 301-line CLAUDE.md | 5 of 5 | $0.239 ($0.221 to $0.256) | 256,450 | 40,066 | 769 | 8 | 17 |
All ten runs passed. The five no-CLAUDE.md runs cost $0.161 to $0.163, and the five CLAUDE.md runs cost $0.221 to $0.256, so the two ranges do not overlap.
Reveneau's arithmetic
The median cost rose by $0.078, from $0.161 to $0.239. That is 48 percent of the no-CLAUDE.md median. Cache write tokens rose by 14,487 per run, from 25,579 to 40,066. Cache read tokens rose by 71,678, from 184,772 to 256,450. Output tokens rose from 659 to 769, and the median turn count from 7 to 8. All of these are differences between medians of five runs, computed by Reveneau from the figures in the table.
The median wall-clock time (the time from start to finish on an ordinary clock) was lower with the file, 17 seconds against 21 without. Reveneau recorded the times and did not investigate this, so this page gives no reason for it.
What changed, by Anthropic's documentation
Anthropic's prompt caching page, read on 4 October 2026, says the model keeps nothing between requests, so Claude Code re-sends the full context each time: the system prompt (Claude Code's own fixed instructions), the project context, every earlier message and tool result, and the new message. The service matches the start of each request against content it recently processed and reads the matching part from the cache instead of processing it again [3]. The same page places CLAUDE.md in the project context layer, which is loaded when the session starts and is included in every request after that [3]. The memory page says CLAUDE.md files in the working directory and the directories above it are loaded at launch, and that a file over 200 lines uses more context [2].
Reveneau's reading of the two token rows follows from that. The 62,746-character file is written to the cache once per run, which is the rise of 14,487 cache write tokens, and it is read back from the cache on every later request in the run, which is the rise of 71,678 cache read tokens over a median of 8 turns. Anthropic's pricing page lists claude-sonnet-4-6 at $6 per million tokens for a one-hour cache write and $0.30 per million for a cache read, against $3 per million for new input, read on 4 October 2026 [4]. By those list prices a one-hour cache write costs more per token than new input and a cache read costs one tenth of new input, which the pricing page states as a cache hit costing 10 percent of the standard input price [4]. Across a run the file is written once and read on every later request, and it costs something on every one of the 8 turns.
The experiment did not separate the two effects by price, because the medians in the table are medians of different columns across the five runs, and the parts do not add up to the cost median exactly. The dollar difference is reported as measured, and the token differences are reported as measured.
Anthropic's 200-line guidance, and what 301 against none can show
Anthropic's costs page, read on 4 October 2026, says to aim to keep CLAUDE.md under 200 lines by including only essentials [1]. The memory page says the same: target under 200 lines per file, because longer files consume more context and reduce adherence, and move instructions that matter for only part of the codebase into path-scoped rules that load only when Claude works with matching files [2]. It adds that splitting a file into imports helps organisation but does not reduce context, because imported files also load at launch [2].
This experiment measured one file of 301 lines against no file, and nothing in between. It cannot say what a 200-line file costs, or a 50-line file, or whether the cost rises in a straight line with length. Reveneau's measurement agrees with the direction of Anthropic's advice: text in CLAUDE.md that the task does not need is paid for on every request. It does not measure the length at which the cost starts to matter, and Anthropic's 200 is Anthropic's figure, with the reasoning given on its memory page. The sister page on how long a CLAUDE.md should be covers that guidance and what to move out.
What this experiment cannot show
One invented task of one failing test in a 30-line project is one task. The medians are medians of five, and nothing here is a statistic about Claude Code in general.
The task did not need anything in the file. A CLAUDE.md the task does need could change the number of turns, in either direction, and on this task turns drive the cost because every turn reads the growing conversation from the cache again. The run did not measure that case.
The task took a median of 7 or 8 turns in this experiment. A longer session reads its CLAUDE.md on every one of its turns, so the same file would cost more tokens across a 40-turn session than across an 8-turn one; the run did not measure that either, and the page on what the benchmark cannot show lists the experiments that were planned and not run.
The model was claude-sonnet-4-6, as Claude Code 2.1.118 resolved the sonnet alias on that day; the same version resolved opus to claude-opus-4-7 and haiku to claude-haiku-4-5-20251001, and neither was run in this experiment. The dollar figures are Claude Code's own list-price estimates on a subscription account. They do not show how much of a plan's allowance the runs used, which no Claude Code output reports.
What Reveneau recommends
Reveneau's recommendation, drawn from Anthropic's documentation and consistent with this measurement: keep CLAUDE.md under Anthropic's 200 lines, put only what every session needs in it, and move instructions that apply to one workflow into a skill or a path-scoped rule so they load when they are used [1][2]. On this task, 301 lines of unused rules added $0.078 to a $0.161 run by Claude Code's figure. A reader whose CLAUDE.md is long and whose tasks are short can repeat the measurement on their own project with the do-it-yourself page before deciding what to cut. The token usage guide explains why the project context layer is paid for on every request.
Common questions
What did the CLAUDE.md experiment in the Reveneau benchmark compare?
The experiment compared two settings on the same task: no CLAUDE.md file, and a 301-line, 62,746-character CLAUDE.md of project rules unrelated to the task, with 20 headed sections of 14 numbered rules each. Both settings ran five times on claude-sonnet-4-6 at default effort with Claude Code 2.1.118 on 4 October 2026, in a fresh copy of the folder each time. The file is published as task/CLAUDE-long.md beside the raw data.
How much did the 301-line CLAUDE.md add to the median cost of the benchmark task?
The median cost rose from $0.161 with no CLAUDE.md to $0.239 with the 301-line file, by Claude Code's own list-price figure on a subscription account. By Reveneau's arithmetic that is $0.078 more, or 48 percent of the no-CLAUDE.md median. The five no-CLAUDE.md runs ranged from $0.161 to $0.163 and the five CLAUDE.md runs from $0.221 to $0.256, so no run of one setting overlapped a run of the other.
Why did the 301-line CLAUDE.md raise cache read tokens more than cache write tokens?
Cache write tokens rose by 14,487 per run and cache read tokens by 71,678, by Reveneau's arithmetic on the medians. Reveneau's reading: the file is written to the cache once per run and read back from the cache on every later request, and the median run had 8 turns, so the same text is read several times. Anthropic's prompt caching page, read on 4 October 2026, says Claude Code re-sends the full context on every request and reads the unchanged start from the cache.
Did the 301-line CLAUDE.md change the number of turns in the benchmark?
The median turn count was 8 with the 301-line CLAUDE.md and 7 without it, and the median output token count was 769 against 659. Reveneau did not read the ten transcripts to find out what the extra turn did, so this page does not say why. One task is one task and these are medians of five runs, so a difference of one turn is reported as what was measured and nothing more.
Did the CLAUDE.md experiment test a file of 200 lines?
No. The experiment measured one 301-line file against no file at all, and tested no length in between. Anthropic's documentation, read on 4 October 2026, says to target under 200 lines per CLAUDE.md file because longer files use more context and reduce adherence. The benchmark cannot say what a 200-line file or a 50-line file would have cost on this task; a later dated run could measure those lengths.
Was the content of the 301-line CLAUDE.md relevant to the benchmark task?
No. The file held 20 headed sections of 14 numbered project rules each, none of which concerned the failing search() test, and the task needed nothing in it. That was deliberate: the experiment set out to measure the cost of text that is loaded on every request and never used. A CLAUDE.md that the task does need could change how many turns the run takes, and the run did not measure that case.
Were the benchmark runs with the 301-line CLAUDE.md slower?
No. The median wall-clock time was 17 seconds with the 301-line CLAUDE.md and 21 seconds without it, even though the CLAUDE.md runs had a median of 8 turns against 7. Reveneau recorded the times from its own clock and did not investigate the difference, so this page offers no reason for it.
Does the CLAUDE.md experiment show that a long CLAUDE.md always costs 48 percent more?
No. The 48 percent is Reveneau's arithmetic on the medians of five runs of one invented task that took a median of 7 or 8 turns, with one file of 301 lines. A longer task with more turns reads the file more times; a shorter file adds fewer tokens; a file the task needs could change the turn count in either direction. The figure describes this task on 4 October 2026 with Claude Code 2.1.118 and claude-sonnet-4-6, and nothing wider.
Where can I read the 301-line CLAUDE.md used in the benchmark?
The file is published as task/CLAUDE-long.md at /benchmarks/claude-code-tokens/2026-10-04/, beside runs.json, runs.csv and the scripts. It is 301 lines of text (342 with the blank lines) and 62,746 characters, arranged as 20 headed sections of 14 numbered rules each. The script run.sh copies it into the fresh task folder as CLAUDE.md for the second setting only, so a reader can confirm exactly what Claude Code loaded in those five runs.
Did every run in the CLAUDE.md experiment pass the test?
Yes. All five runs with no CLAUDE.md and all five with the 301-line file passed, graded by re-running node --test test/ledger.test.js outside Claude Code after each run and treating exit code 0 as a pass. The experiment therefore compares the cost of two settings that both did the job; it says nothing about whether a long CLAUDE.md changes the quality of the result, because the task has only one correct outcome.
Why did the benchmark use a CLAUDE.md of rules unrelated to the task?
Because the question was the cost of instructions that are loaded on every request and never used, which Anthropic's costs page, read on 4 October 2026, names as the reason to move workflow-specific instructions out of CLAUDE.md into skills. A file of relevant rules would mix two effects, the cost of the text and any change in how the model works, and five runs could not separate them. Unrelated rules isolate the first effect.
References
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, How Claude remembers your project (code.claude.com), read 4 October 2026
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
- Anthropic, Pricing (platform.claude.com), read 4 October 2026
More in Experiments
The model choice experiment: Haiku, Sonnet and Opus on one task
In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times on each of three models, with no CLAUDE.md and default effort, using Claude Code 2.1.118 on a Claude subscription. The aliases resolved to claude-haiku-4-5-20251001, claude-sonnet-4-6 and claude-opus-4-7. The median costs by Claude Code's own list-price figure were $0.060, $0.161 and $0.489. By Reveneau's arithmetic, Haiku's median was 37 percent of Sonnet's and Opus's was 3.0 times Sonnet's. All 15 runs passed. Haiku took more turns, a median of 9 against 7. One task is one task, the figures are medians of five, and newer models were not measured.
The effort level experiment: low, medium and high on Sonnet 4.6
In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times at each of three effort levels, passed with the --effort flag, on claude-sonnet-4-6 with no CLAUDE.md and Claude Code 2.1.118 on a Claude subscription. The median costs by Claude Code's own list-price figure were $0.150 at low, $0.162 at medium and $0.170 at high. By Reveneau's arithmetic the gap between low and high is $0.020, or 13 percent of the low median, and the ranges overlap: one low run cost $0.208, above every high run. The difference in medians came from turns, 6 against 7; output token medians are within 33 tokens of each other. All 15 runs passed. One task is one task, and the figures are medians of five.
The output-trimming hook experiment: a test that prints 3,000 lines
In Reveneau's Claude Code token benchmark of 4 October 2026, a second test printing 3,000 lines was added and the task ran five times with no hook and five times through Anthropic's example PreToolUse hook that filters test output, on claude-sonnet-4-6 with Claude Code 2.1.118 on a Claude subscription. The hook runs cost more: a median of $0.201 against $0.159 by Claude Code's own list-price figure, with a median of 10 turns against 7. The transcripts show why: in every direct run the model itself added | tail -20 to the test command, keeping the last 20 lines, and in the hook runs the filter left no output on a pass, so the model ran the tests again. One task is one task; the figures are medians of five.