AI NewsModels & agentsReported

GPT-6 Sol answered 12 of 15 developer tasks correctly, Claude Opus 5.5 answered 15 of 15, for six times the cost

An independent test that ran three developer tasks five times each through GPT-6 Sol and Claude Opus 5.5 found Sol was faster and cheaper on every task, but missed 3 of 15 runs while Opus 5.5 was perfect on all 15.

AI News

Editorial2 min read

LinkedInX

Why it mattersSol earns its price on work a human or a test suite will check, and Opus 5.5 earns its price on work nobody re-reads before it reaches a customer.

A single-run comparison of two frontier models can hide the difference that actually shows up in production, which is whether the same prompt gives the same answer the second and third times you send it. Jessica Wachtel at The New Stack ran GPT-6 Sol and Claude Opus 5.5 through three developer tasks five times each on 28 September 2026, and Opus 5.5 was correct on all 15 runs while Sol was correct on 12. Sol's total cost across those 15 runs was $2.68, against $16.72 for Opus 5.5.

Wachtel called both models through their vendor APIs with identical prompts, at the highest effort each model exposes. She logged tokens, cost at list price, and time for every call. OpenAI lists Sol at $2 per million input tokens and $10 per million output tokens; Anthropic lists Opus 5.5 at $4 and $20.

The three tasks

The first task was triaging 40 failed CI jobs. Both models were perfect on all five runs. Opus 5.5 averaged 1 minute 27 seconds and 11,127 output tokens per run, at $0.24. Sol averaged 18 seconds and 1,143 output tokens, at $0.02. The New Stack notes Sol cost 8 percent of what Opus 5.5 did on this test, close to the 9 percent cheaper number that OpenAI's own AutomationBench claims for Sol against Opus 5.

The second task was reading a 3,664-line outage postmortem. Opus 5.5 was perfect on all five runs; Sol was perfect on two. Two of Sol's misses left the same customer off a refund list, whose original charge was confirmed 52 seconds after the retry had already succeeded. Another run miscounted failed checkouts as 27 instead of 28. Opus 5.5 averaged 6 minutes 44 seconds and $1.68 per run; Sol averaged 1 minute 41 seconds and $0.30, and read the file in 25 percent fewer input tokens.

The third task was implementing a spec graded by 120 hidden tests. Opus 5.5 passed every test on every run. Sol passed all 120 tests on four runs; on the fifth, it left a stray closing parenthesis on line 78 and the module crashed on import.

The pattern in the misses

Sol won on every axis except accuracy. The New Stack writes that Sol's misses were "the kind that slip past review": a customer left off a refund list, a file that will not import. On the surface each is one wrong number or one wrong character; in a queue it is a customer nobody credits and a service that will not deploy. Opus 5.5 cost about six times what Sol did overall, which is the number a team weighs against the runs a person will not re-read.

Wachtel's recommendation: use Sol for high-volume work where a person or a test suite checks the output, and where you can afford to run it twice at those prices. Use Opus 5.5 when a wrong answer is expensive and nobody is checking, such as a reconciliation or an incident report that goes straight to a customer. The choice turns on what a team already does with the output, since the sticker price only matters when the answer is trusted.

Source

GPT-6 Sol vs. Claude Opus 5.5: Is cheaper important when results aren't consistent? by Jessica Wachtel, The New Stack, 28 September 2026.

Reported byThe New Stack

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX