The output-trimming hook experiment: a test that prints 3,000 lines
In Reveneau's Claude Code token benchmark of 4 October 2026, a second test printing 3,000 lines was added and the task ran five times with no hook and five times through Anthropic's example PreToolUse hook that filters test output, on claude-sonnet-4-6 with Claude Code 2.1.118 on a Claude subscription. The hook runs cost more: a median of $0.201 against $0.159 by Claude Code's own list-price figure, with a median of 10 turns against 7. The transcripts show why: in every direct run the model itself added | tail -20 to the test command, keeping the last 20 lines, and in the hook runs the filter left no output on a pass, so the model ran the tests again. One task is one task; the figures are medians of five.
Published October 4, 2026. Editorial.
Key takeaways
- Through the trimming hook the median cost was $0.201 against $0.159 without it, by Claude Code's own list-price figure on a subscription account, with a median of 10 turns against 7; all ten runs passed.
- In every one of the five direct runs the model appended | tail -20 to the test command on its own, so the 3,000 printed lines never entered the conversation and the largest tool result was 869 characters.
- In the five hook runs the filtered command produced no output when the tests passed, because the pattern matches only failure words and Node's summary lines contain none; the model saw only a 31-character notice of that and then ran between 2 and 8 test commands per run against 1 in the direct runs.
- Anthropic's documentation describes the hook with an example of a 10,000-line log file; that is Anthropic's illustration, and this experiment did not test it.
- Reveneau's reading: a filter that removes the test runner's success summary costs turns, so a filter should keep the summary line.
This is the fourth experiment in Reveneau's Claude Code token benchmark. It asks whether a hook that filters test output lowers the cost of a run whose tests print a lot. Reveneau is an AI software development consultancy whose code is written by AI and whose every change must pass an eval suite (a set of automated tests) before release, so test output enters its conversations many times a day, and a filter that works would be worth having in every build. Reveneau is independent of Anthropic. The full method is on the method page.
The words used on this page
A token is a piece of text the model reads or writes. The context window is all the text the model reads in one request, and the output of every command the model runs goes into it. Claude Code sends the whole conversation again on every request; the service stores the unchanged start, and tokens stored this way are cache write tokens, while tokens read back from the store on a later request are cache read tokens [5]. CLAUDE.md is a file of instructions loaded at the start of every session; this experiment used none. The effort level is a setting for how much reasoning the model does; this experiment passed no flag for it. A hook is a shell command that Claude Code runs on its own at a fixed point in its work: Anthropic's hooks guide, read on 4 October 2026, says hooks are user-defined shell commands that Claude Code runs at specific points, which gives deterministic control, meaning the action always happens instead of depending on the model choosing to do it [3]. A non-interactive run is claude -p with a prompt: Claude Code does the task and exits. A turn is one request to the model and its reply. A median is the middle value of five sorted runs.
What was compared
A second test file was added to the task. It prints 3,000 lines, each 67 to 73 characters long, and passes. It is published as task/noisy.test.js. The published files call it the noisy test; this page calls it the test that prints 3,000 lines.
The prompt was changed so that the model would run both test files with one command and see the long output:
One test in test/ledger.test.js fails. Fix the code in src/ so that
node --test test/ledger.test.js test/noisy.test.jspasses (run that exact command to check; the noisy test prints a lot). Change only src/. Do not change the tests. When done, stop.
The two settings were "direct", with no hook, and "hook", with a filter attached. Both ran five times on claude-sonnet-4-6 with no CLAUDE.md, no --effort flag, Claude Code 2.1.118 on a Claude subscription, the flags on the method page, and a fresh copy of the folder for every run, on 4 October 2026.
The hook
The hook is Anthropic's own example. Its costs page, read on 4 October 2026, has a section on offloading processing to hooks, which says a hook can preprocess data before Claude sees it, and gives an illustration: instead of Claude reading a 10,000-line log file to find errors, a hook can grep for ERROR and return only matching lines, reducing context from tens of thousands of tokens to hundreds [1]. That is Anthropic's illustration, and this experiment did not test it. The same page then gives a PreToolUse hook that filters test output to show only failures: a settings entry with the matcher Bash (the rule that says which tool the hook applies to), and a shell script that reads the command, checks whether it starts with npm test, pytest or go test, and if so rewrites it to pipe through grep -A 5 -E '(FAIL|ERROR|error:)' | head -100 [1].
Anthropic's hooks reference, read the same day, explains the two fields the script returns. PreToolUse fires before a tool call executes and can control whether it proceeds [2]. permissionDecision: "allow" skips the permission prompt, and updatedInput modifies the tool's input before execution, replacing the entire input object, so the script returns the whole original input with only the command changed [2]. The reference adds that Claude Code evaluates permission rules against the input the hook returns [2]. In plain words: every time the model asks to run a shell command, the script sees the command first, and if it is a test command the script swaps in a filtered version before anything runs.
Two strings in the published filter-test-output.sh differ from Anthropic's example. The test-runner pattern is ^node --test in place of ^(npm test|pytest|go test), because the task runs its tests with node, as the comment at the top of the script says. The word list is (FAIL|ERROR|error:|not ok), with not ok added; the script records no reason for that second change. The rest is as Anthropic published it. The hook was attached with the --settings flag, which Anthropic's CLI reference describes as a settings file whose values override the same keys for that session [4]; the settings file registers the script under PreToolUse with the matcher Bash.
The result
Every figure is Reveneau's own measurement from the run of 4 October 2026, a median of five, with the dollar figure being Claude Code's own list-price figure on a subscription account. The rows are in runs.json.
| Setting | Pass | Cost median (min to max) | Cache read median | Cache write median | Output median | Turns median | Seconds median |
|---|---|---|---|---|---|---|---|
| direct, no hook | 5 of 5 | $0.159 ($0.149 to $0.173) | 183,709 | 25,154 | 660 | 7 | 16 |
| through the trimming hook | 5 of 5 | $0.201 ($0.164 to $0.245) | 290,134 | 26,301 | 984 | 10 | 21 |
All ten runs passed. The hook setting cost more: $0.042 more at the median by Reveneau's arithmetic, with a median of 10 turns against 7, 290,134 cache read tokens against 183,709, and 984 output tokens against 660. Its range, $0.164 to $0.245, was the widest of any setting in the benchmark apart from Opus.
What the transcripts show
Claude Code saves a transcript per session, and Reveneau read the transcripts of all ten runs to find out why the filtered setting cost more. Two things were found.
In every one of the five direct runs, the model itself appended | tail -20 to the test command, so only the last 20 lines of output entered the conversation and the 3,000 printed lines never did. The largest tool result in those five runs was 869 characters. The direct setting therefore had no long tool result to filter: its median of 183,709 cache read tokens and 7 turns compares with 184,772 and 7 in the no-CLAUDE.md Sonnet setting on the method page, which had no long output at all.
In the five hook runs, the filtered command produced no output when the tests passed, and the model saw only Claude Code's 31-character notice, (Bash completed with no output). The pattern matches only failure words, and Node's summary lines, # tests 5, # pass 5 and # fail 0, contain none of them. The model then ran the test command again in other forms to confirm the result: with | head -60, with | grep, with ; echo "exit: $?". Across the five hook runs there were between 2 and 8 test commands per run, against 1 in each direct run. One hook run reached 15 turns. In all five hook runs the last test command ended with ; echo and the exit code; the hook rewrote that command too, placing the filter after the echo, so the filter applied to the echo alone, the test's full output came back, and Claude Code saved it to a file and showed the model a preview of its first part.
Reveneau's reading
A filter that removes the success summary costs turns. The model needs to know whether the tests passed, and when the filter leaves it nothing to read, it asks again, and every extra turn re-sends the whole conversation and reads the growing conversation from the cache once more, which is where the extra 106,425 cache read tokens at the median come from by Reveneau's arithmetic. On this task the model's own habit of adding | tail -20 already kept the output small, so there was no large tool result for the filter to remove.
Two conclusions follow, both limited to this task. First, a filter should keep the test runner's summary line, so the model learns the result in one read. Reveneau did not run a hook that does this, so there is no measurement of one yet; it is a candidate for a later dated run. Second, before adding a filter, read a few transcripts to see what the model already does with long output, because on this task it had solved the problem itself.
Anthropic's 10,000-line log illustration describes a different situation: a file the model would read in full, with no habit of trimming it. That case was not measured here and nothing on this page says the hook would fail there. The sister page on hooks that trim output covers Anthropic's guidance on where such a hook helps, and the agent costs guide covers the other way Anthropic suggests for long output, which is to give it to a subagent, a separate conversation the main one starts for part of the work.
What this experiment cannot show
One invented task of one failing test in a 30-line project, with one added test that prints 3,000 lines, is one task. The medians are medians of five, and nothing here is a statistic about Claude Code in general or about hooks in general.
The experiment ran one filter, Anthropic's example with two strings changed, and did not run a version that keeps the summary line. It did not measure a log file the model would read in full, which is Anthropic's own example. It did not measure what a filter does when a test fails, because every run passed, although the fail case is the one where the filter's pattern would match lines and return them.
The model was claude-sonnet-4-6 as resolved by Claude Code 2.1.118 on 4 October 2026; the same version resolved opus to claude-opus-4-7 and haiku to claude-haiku-4-5-20251001, and neither was run with the hook. Another model could handle long output differently; the tail -20 habit was observed in the ten Sonnet runs of this experiment and was not looked for on another model. The dollar figures are Claude Code's own list-price estimates on a subscription account. The limits page lists everything the run cannot show.
What Reveneau recommends
Reveneau's recommendation, drawn from Anthropic's documentation and consistent with this measurement: if you add an output filter, make it keep the line that states the result, and check a transcript afterwards to confirm the model reads that line once and moves on. Anthropic's costs page says to verify a hook with /hooks and that a debug file records when a hook rewrites a command [1]; a transcript shows what the model did next. A reader can measure their own filter against no filter with the do-it-yourself page, and the token usage guide explains why each extra turn costs a full read of the conversation.
Common questions
What did the hook experiment in the Reveneau benchmark measure?
It measured the same task with a second test file added that prints 3,000 lines and passes, run five times with no hook and five times through Anthropic's example PreToolUse hook from its costs page, adapted to Node's test runner. Both settings ran on claude-sonnet-4-6 with no CLAUDE.md and Claude Code 2.1.118 on 4 October 2026, in a fresh copy of the folder each time. The question was whether the filter lowered the cost of a run that produces long test output.
Why did the runs through the trimming hook cost more on the benchmark task?
Because they took more turns: a median of 10 against 7, with a median cost of $0.201 against $0.159 by Claude Code's own list-price figure on a subscription account. Reveneau read all ten transcripts. In the hook runs the filtered test command produced no output on a pass, since the pattern matches only failure words and Node's summary lines contain none, so the model ran the tests again in other forms to confirm the result, between 2 and 8 test commands per run against 1 in the direct runs.
What did the direct runs do with the 3,000 lines of test output?
In every one of the five direct runs, the model itself appended | tail -20 to the test command, so only the last 20 lines of output entered the conversation and the 3,000 printed lines never did. The largest tool result in those five runs was 869 characters. This is why the direct setting, with no filter at all, had a median of 183,709 cache read tokens and 7 turns, against 184,772 and 7 in the no-CLAUDE.md Sonnet setting, which had no long output at all.
How was Anthropic's example hook changed for the benchmark?
Two strings differ from the example on Anthropic's costs page, read on 4 October 2026. The test-runner pattern was changed from ^(npm test|pytest|go test) to ^node --test, because the task runs its tests with node, as the comment in the published script says. The word list in the grep command gained not ok, so it reads (FAIL|ERROR|error:|not ok). The rest, including grep -A 5, head -100, permissionDecision allow and updatedInput, is as Anthropic published it.
What did the filtered test command return when the tests passed?
Nothing. The filter keeps only lines that match FAIL, ERROR, error: or not ok, plus five lines after each. Node's summary lines, # tests 5, # pass 5 and # fail 0, contain none of those words, so on a passing run no line came back and the model saw only Claude Code's 31-character notice of an empty result. It then ran the test command again with | head -60, with | grep, or with ; echo "exit: $?" to learn whether the tests had passed.
Does the hook experiment contradict Anthropic's 10,000-line log example?
No. Anthropic's costs page, read on 4 October 2026, illustrates the hook with a 10,000-line log file in which a grep for ERROR returns only matching lines, reducing context from tens of thousands of tokens to hundreds. That is Anthropic's illustration and the experiment did not test it. The experiment measured one task where the model already piped its output through tail and where the filter removed the success summary; on that task, the hook cost more.
How many turns did the longest hook run take?
15 turns. The five hook runs had a median of 10 turns, against 7 for the direct runs, and the longest reached 15, with every run still under the --max-turns 25 cap. The hook runs' cost ranged from $0.164 to $0.245 by Claude Code's figure, the widest range of any setting apart from Opus, and the median cache read count was 290,134 against 183,709 in the direct runs, because each extra turn re-reads the growing conversation from the cache.
What does Reveneau conclude about filters that remove the success summary?
That on this task a filter which removes the test runner's summary line costs turns, because the model runs the tests again to learn the result, and that a filter should keep the summary line. This is Reveneau's reading of ten transcripts of one task on 4 October 2026, with Claude Code 2.1.118 and claude-sonnet-4-6. Reveneau did not run a version of the hook that keeps the summary line, so it has no measurement of one yet.
Is the hook script from the benchmark published?
Yes. The script is at /benchmarks/claude-code-tokens/2026-10-04/scripts/filter-test-output.sh beside the test file task/noisy.test.js that prints 3,000 lines, the run script run2.sh, and the raw data in runs.json. The script is Anthropic's example with the two changes described on this page. A reader can see the exact command the hook rewrote each test run into, and the exact word list that left no output on a pass.
Why did the hook experiment use a different prompt from the other experiments?
Because the task had two test files instead of one. The prompt named both files in one exact command, node --test test/ledger.test.js test/noisy.test.js, asked the model to run that command to check, and said that the second test prints a lot. Without that instruction the model could have run only the failing test file and never produced the long output the experiment set out to measure. The prompt is in the published run2.sh.
How was the hook attached to the benchmark runs?
With the --settings flag, which Anthropic's CLI reference, read on 4 October 2026, describes as a settings file whose values override the same keys for that session. The settings file registers the script as a PreToolUse hook with the matcher Bash, so it runs before every shell command. Anthropic's hooks reference says a PreToolUse hook can return permissionDecision allow with updatedInput to replace the command before it runs, which is what the script does for commands starting with node --test.
Did every run in the hook experiment pass?
Yes. All five direct runs and all five hook runs passed, graded by re-running node --test test/ledger.test.js outside Claude Code after each run and treating exit code 0 as a pass. Every hook run fixed the failing test; the extra turns in the hook setting came from confirming a pass, since the filter left no summary line to read. The experiment compares the cost of two settings that both did the job.
References
- Anthropic, Manage costs effectively (code.claude.com), read 4 October 2026
- Anthropic, Hooks reference (code.claude.com), read 4 October 2026
- Anthropic, Automate actions with hooks (code.claude.com), read 4 October 2026
- Anthropic, CLI reference (code.claude.com), read 4 October 2026
- Anthropic, How Claude Code uses prompt caching (code.claude.com), read 4 October 2026
More in Experiments
The CLAUDE.md length experiment: 301 lines against none
In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times with no CLAUDE.md and five times with a 301-line CLAUDE.md of project rules unrelated to the task (62,746 characters; 342 lines when the blank lines between sections are counted), on claude-sonnet-4-6 at default effort with Claude Code 2.1.118. The median cost rose from $0.161 to $0.239 by Claude Code's own list-price figure on a subscription account, which is $0.078 or 48 percent more by Reveneau's arithmetic. Cache write tokens rose by 14,487 and cache read tokens by 71,678. All ten runs passed. One task is one task, the figures are medians of five, and the experiment measured one file against none, with nothing in between.
The model choice experiment: Haiku, Sonnet and Opus on one task
In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times on each of three models, with no CLAUDE.md and default effort, using Claude Code 2.1.118 on a Claude subscription. The aliases resolved to claude-haiku-4-5-20251001, claude-sonnet-4-6 and claude-opus-4-7. The median costs by Claude Code's own list-price figure were $0.060, $0.161 and $0.489. By Reveneau's arithmetic, Haiku's median was 37 percent of Sonnet's and Opus's was 3.0 times Sonnet's. All 15 runs passed. Haiku took more turns, a median of 9 against 7. One task is one task, the figures are medians of five, and newer models were not measured.
The effort level experiment: low, medium and high on Sonnet 4.6
In Reveneau's Claude Code token benchmark of 4 October 2026, one small coding task was run five times at each of three effort levels, passed with the --effort flag, on claude-sonnet-4-6 with no CLAUDE.md and Claude Code 2.1.118 on a Claude subscription. The median costs by Claude Code's own list-price figure were $0.150 at low, $0.162 at medium and $0.170 at high. By Reveneau's arithmetic the gap between low and high is $0.020, or 13 percent of the low median, and the ranges overlap: one low run cost $0.208, above every high run. The difference in medians came from turns, 6 against 7; output token medians are within 33 tokens of each other. All 15 runs passed. One task is one task, and the figures are medians of five.