Effortless auto-picks Claude Code's reasoning effort for each prompt
Effortless is a new MIT-licensed Claude Code plugin that reads each prompt with Haiku 5.5 and picks a reasoning effort before Opus runs it, and the author reports a 21% cost cut against Anthropic's recommended medium effort on 96 runs of four coding tasks.
Image: GitHub
Why it mattersA team paying for Opus on Claude Code pays for every token the model thinks with, and most prompts in a working day do not need the highest effort. A judge that picks per prompt turns that overhead into a setting that nobody has to touch.
Opus on Claude Code runs at Anthropic's recommended medium reasoning effort by default, and a user who wants a cheaper or deeper answer flips the control by hand. Effortless, a Claude Code plugin published by the developer HeyCubit, reads each prompt with Haiku 5.5 and sets that control for every message on its own. The repository is MIT-licensed and has gained 100 stars in four days.
The author also published a benchmark. On four coding tasks, 12 answers each at medium, high and xhigh with no plugin and 12 answers with the plugin on, through Anthropic's API, Opus 5.5 cost $0.190 per task at medium, $0.209 at high, $0.257 at xhigh, and $0.149 with Effortless, with every one of the 48 answers scored correct. That is 21% less than medium, 29% less than high and 42% less than xhigh. On Sonnet the plugin came out even against medium, saved 10% against high and 35% against xhigh. The author names the limits in the same README: four tasks, three runs each, "mostly easy work where effortless goes low", on fresh chats through the API rather than long chats on a plan.
Where the saving comes from
The plugin's note on its own numbers is more useful than the headline. Across 80,000 requests of real Claude Code use the author measured, 76% of the cost was each request re-reading the whole chat, 16% was cache writes and 8% was the model's output. A lower effort does not shorten any one reply by much; it shortens the number of request-reply loops the model takes to finish, and every loop re-reads the chat. Picking low effort on a prompt that does not need it saves on the 76%.
A bar above the prompt also shows how full the current chat is, how long the prompt cache will stay warm, and turns a Compact button into a Handoff button when Haiku judges that a fresh chat would be cheaper to carry on in. Compactions, both manual and the automatic one Claude Code runs, are written by Haiku 5.5 by default, which Anthropic recommends for compaction at a lower price than the chat model.
How the judge is set up
Haiku 5.5 runs on the user's own Claude login and needs no key. The author reports a median judge time of 0.75 seconds, and a short follow-up like "ok" or "go" keeps the previous effort and asks no judge. The author offers a second judge, TypeSafe's Jev decision model, at a reported 0.24 seconds a call behind an API key, with Haiku stepping in whenever Jev is unsure. The judges' accuracy is reported from a 73-case test suite: Haiku was right on 92 to 93% of all cases and on 95% of a 20-case held-out set that was never used to tune the prompts, and Jev was right on 99% and 100% of the same two sets. The author states that the 53 non-held-out cases were used to tune the judge prompts and that 73 cases is a small set.
The plugin does not switch to a cheaper model mid-chat on its own: the author's measurement on an 80,000-token chat shows that writing the chat into a second model's cache costs five times what staying on a warm Opus does, and ten times for Sonnet, so the model-switch option is marked experimental and off by default. Effort only changes on Opus 5.5 and Sonnet 5.5, since older models rewrite most of the prompt cache when effort changes between requests.
Source
Primary source: HeyCubit/effortless on GitHub.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.