The New Stack ran five Terminal-Bench-Science tasks on Claude Fable 5.1 and Fable 5 with a $12 cap per task, and Fable 5.1 solved one, Fable 5 solved none

Image: The New Stack
Why it mattersAnthropic's headline is that Fable 5.1 more than doubles Fable 5 on a science benchmark. On a laptop-sized budget the gap collapses, so the harness and the money you give it matter as much as which model you picked.
Jessica Wachtel at The New Stack picked five tasks from the Terminal-Bench-Science benchmark, one from each science field, and ran both Claude Fable 5 and Fable 5.1 on each under a $12 cost cap and a 60-turn limit. Anthropic reports 24.7% for Fable 5 and 52.6% for Fable 5.1 on the full 70-task suite. The New Stack's smaller run got 0% for Fable 5 and 20% for Fable 5.1.
Anthropic has not disclosed the harness or budget behind its numbers. Terminal-Bench-Science allows each model up to eight hours per task, and when the benchmark's own leaderboard tested Fable 5 through Claude Code at maximum effort, it spent $14,180 across 210 attempts, which works out to $67.52 per attempt. The New Stack's runs used a plain terminal, $12 per task, and 60 turns.
The five tasks and what each model did
The one task Fable 5.1 solved was symbolic regression. It found a hidden formula behind a yes-or-no label, stopped on its own after 27 turns and 11.8 minutes, and cost $1.96 for 27,088 output tokens. Fable 5 used all 60 turns, ran for 53.5 minutes across two attempts, spent $4.20 and then $6.38, and failed both times.
Lorenz-96 assimilation, an atmospheric-model reconstruction task, defeated both models. Fable 5 hit the $12 cap at 45 turns and $12.63. Fable 5.1 used all 60 turns for 126 minutes and $10.70. On the leaderboard, Earth sciences is also where Fable 5 scores close to zero, so both failing here is consistent.
Reactor safety control asked for a controller that never exceeds a temperature limit. Neither model wrote one that passed. Fable 5.1 hit the turn limit with 157,710 output tokens and $11.53; Fable 5 hit the cost cap at $12.04. In the foraging cognitive model task, Fable 5.1 declared itself finished at 43 turns and $5.65, but the grader rejected its answer. Nanoindentation, extracting material properties from raw indentation curves, was also failed by both.
What the totals say
Across the five tasks, Fable 5 solved 0, spent $53.59, ran 388 minutes, and hit the cost cap four times. Fable 5.1 solved 1, spent $40.75, ran 262 minutes, and never hit the cost cap. Output token totals were similar: 430,356 for Fable 5 and 455,792 for Fable 5.1.
Wachtel is careful about what five tasks can and cannot show. She writes that getting these results by chance is plausible even if Anthropic's published scores are exactly right, so this run leaves the doubling claim open.
What did show up was on the bill. Fable 5.1 failed faster and cheaper than Fable 5. For a team choosing between the two on their own budget rather than Anthropic's, the numbers point to a difference in spend and time to failure. Accuracy on this small sample looked the same. Wachtel writes that to approach the leaderboard-scale numbers, the harness, the hours per task, and a larger budget matter as much as which model runs the task.
Source
The New Stack: Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet
Reported by: The New Stack
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually shipping with them. Short, and only when there is something worth reading.


