AI NewsModels & agentsReported

GPT-6.1 Sol tied Astra on 15 engineering runs at one-fifth the cost

Jessica Wachtel at The New Stack ran GPT-6.1 Sol and GPT-6 Astra through three real engineering tasks, five runs each, and both models got every answer right on all 15 runs. Sol cost $2.66 to Astra's $14.77 in total and was faster on the two longer tasks.

AI News

Editorial2 min read

LinkedInX

Why it mattersA team paying for Astra for CI triage, incident log questions and spec-driven coding has one run of real tests showing Sol clears every bar they were paying extra for.

The question a team pays Astra for is whether the extra money buys answers Sol would get wrong. Jessica Wachtel at The New Stack ran both models through the same three real engineering tasks she had used on earlier models, five runs each, and both finished 15 of 15 runs with every answer correct. Sol cost $2.66 for the whole set to Astra's $14.77.

How the tests were run

Wachtel called both models through the OpenAI Responses API with identical prompts, reasoning effort set to max and a 64,000-token output limit. The three tasks each map to one of OpenAI's own marketing claims for the two models: a CI triage task of 40 calls per run, an incident-logs task of seven questions over a 113,966-token input, and a resolver spec with 120 hidden tests the output has to pass. She did not publish the prompts, because they use multiple files and repos rather than copy-paste snippets.

What the three tests showed

On CI triage, both models made all 40 calls correctly on all five runs. Astra averaged 24 seconds per run against Sol's 31 seconds, so Astra was faster, and Sol cost $0.02 per run against Astra's $0.11.

On incident logs, Sol averaged 2 minutes 20 seconds against Astra's 2 minutes 53 seconds. Both read 113,966 input tokens. Sol cost $0.31 per run against Astra's $1.57, and both answered all seven questions correctly every time.

On the resolver spec, both models passed all 120 hidden tests on every run. Sol averaged 7 minutes 7 seconds against Astra's 10 minutes 6 seconds. Astra wrote 25,207 output tokens per run to Sol's 19,637. The cost gap was the largest: $0.20 for Sol against $1.28 for Astra.

Across the 15 runs, Sol took 49 minutes 54 seconds for $2.66 and Astra took 1 hour 7 minutes for $14.77. Sol tied Astra on accuracy and won on speed in two of the three tasks.

What this does not show

Wachtel is clear that three tasks cannot prove the two models are equal on everything, and Artificial Analysis' index still puts Astra ahead at max effort. The tasks were chosen because they map to OpenAI's own claims for the models, so a workload outside those shapes could land differently. Wachtel ran the same tests on GPT-6 Sol, so a reader can compare GPT-6.1 Sol's 15 of 15 to that earlier run: GPT-6 Sol missed customers on two incident-log runs and crashed the resolver with a stray parenthesis on one run, which GPT-6.1 Sol and Astra both now pass.

The practical shape of the finding is narrow and useful. A team already paying for Astra on tasks that look like these three has a reason to try Sol with max reasoning effort on the same prompts, and a baseline cost number to compare against: five runs of each real task, same settings, same prompts. The resolver spec saved $1.08 per run against identical passing output, so a team that runs an agent on a spec file a few hundred times a week has the saving measured.

Source

Primary source: GPT-6.1 Sol vs. GPT-6 Astra: Same accuracy at 18% of the cost, Jessica Wachtel at The New Stack.

Reported byThe New Stack

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX