AI NewsModels & agentsReported

Sonnet 5.5 beat Opus 5.5 on cost and accuracy in 15-run test

Jessica Wachtel at The New Stack ran Sonnet 5.5 and Opus 5.5 through three hidden-grader coding tests five times each, and Sonnet was perfect on 15 of 15 runs for 42 percent less money.

AI News

Editorial2 min read

LinkedInX

Why it mattersThe newer and cheaper model in Anthropic's pair also finished more work correctly, so the default for hard coding tasks moves to Sonnet 5.5 and Opus 5.5 is kept for agent loops.

The newer and cheaper model in Anthropic's pair also turned out to be the more accurate one. Jessica Wachtel at The New Stack ran Sonnet 5.5 and Opus 5.5 against three hidden coding tests, five runs per test per model, and reports that Sonnet 5.5 was perfect on 15 of 15 runs for a total of $12.69, while Opus 5.5 was perfect on 13 of 15 for $22.07.

The three tests were an agentic bug fix on a small Python repo with four planted bugs, a dependency resolver written from a two-page spec, and three concurrency bugs in an asyncio job queue. Each test was graded against a hidden suite the models never saw. Both models called the Anthropic API with identical prompts, adaptive thinking and the maximum effort setting.

Where each model won

On the agentic bug fix, Opus 5.5 finished about 35 percent faster than Sonnet 5.5. Sonnet used twice as many tokens for the same answer and hit the 32,000-token output limit per step on four of its five runs on the first try. Wachtel raised the limit to 128,000 and the reruns passed, at an extra cost of about $1.40 that she folds into Sonnet's total as $0.98 per run versus Opus 5.5's $0.75.

Sonnet 5.5 won the resolver test: both models passed all 120 hidden tests on all five runs, and Sonnet averaged $0.82 per run against Opus 5.5's $1.42. The concurrency test was Sonnet's clearest win. It fixed every race condition on every run, while Opus 5.5 passed three of five and spent all 128,000 output tokens thinking without producing an answer on the other two.

What Artificial Analysis found, and what Wachtel found

Anthropic sets Sonnet 5.5 at $2 per million input tokens and $10 per million output tokens, half the per-token price of Opus 5.5 at $4 and $20. Half the price per token did not deliver half the bill in these tests. Sonnet 5.5 often used more tokens to finish, so the saving came in at about 36 percent across the 15 runs rather than 50 percent. Artificial Analysis, cited by Wachtel, had reported Sonnet 5.5 costing more per task than Opus 5.5 at max effort and scoring lower on its Intelligence Index. Wachtel's three tests went the other way on both.

The usable rules she draws from the result: raise Sonnet 5.5's output limit before using it on hard work, because it thinks longer in a single step than Opus 5.5 does; and keep Opus 5.5 for agent loops, where the agentic test showed it finishing faster at lower total cost once Sonnet's failed runs are counted.

A model-choice rule tuned to a benchmark score does not survive one real test, and this one came from the same writer who ran the earlier Opus 5.5 vs Opus 5 and Opus 5.5 vs Fable 5.1 pieces, using the same hidden-grader method. The three results together say that cost and accuracy for a specific pair of models are a job for a few tests on your own code, not a scoreboard reading.

Source

Reported byThe New Stack

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX