AI NewsModels & agentsReported

Artificial Analysis says GPT-6.1 Sol scores one Intelligence Index point below GPT-6 Astra at 72 cents a task against 3.26 dollars

Artificial Analysis published its benchmark run of GPT-6.1 Sol on 29 September 2026 and reported it one Intelligence Index point below GPT-6 Astra, at $0.72 per task against $3.26 for Astra, and up 12 points on Terminal-Bench 4.0 versus the GPT-6 Sol it replaced after seven days.

AI News

Editorial2 min read

LinkedInX

Why it mattersAn independent run gives teams the second-source number they could not get from OpenAI's own launch, and the cost gap is wide enough to make Sol the default for most coding and agent work with Astra kept for the harder tickets.

A smaller, cheaper model is a routing choice a team has to defend with numbers nobody at the vendor collected, and OpenAI shipped GPT-6.1 Sol last week without that second source. Artificial Analysis published its own benchmark run of Sol on 29 September 2026, and the numbers in the report give a team the independent read the launch day did not: one Intelligence Index point below GPT-6 Astra, at $0.72 per task against $3.26 for Astra on the same measure.

What the report says changed in seven days

GPT-6.1 Sol replaced GPT-6 Sol on 29 September, seven days after the earlier model shipped, and Artificial Analysis reports a four-point gain on its Intelligence Index over that version. The run also places GPT-6.1 Sol one point below GPT-6 Astra on the same index, so OpenAI's smaller model is now scoring at the edge of its flagship. Terminal-Bench 4.0, which Artificial Analysis says is its coding-agent harness benchmark, moved up 12 points for the newer Sol. The Coding Agent Index at max reasoning effort is 3 points higher than the one-week-old Sol and 2 points below Astra. Named individual benchmarks that moved: AA-Briefcase v1.1 up roughly 4 points, GDPval-AA v2.1 up roughly 5, Humanity's Last Exam up 5, and GDP.pdf up 6.

The honest number inside the headline number

Artificial Analysis reports that AA-Omniscience accuracy at max reasoning effort rose 8 points on the newer Sol, and that the hallucination rate measured on the same test dropped from 60 percent to 54 percent. The hallucination number is the part a reader has to carry forward: 54 percent means that when GPT-6.1 Sol is wrong on this test, it makes up an answer in 54 cases out of 100. The six-point drop from 60 is a real move, and 54 percent is still a long way from a model that reliably says it does not know.

Where the price sits

OpenAI's listed per-million-token price is $2 for input and $10 for output, the same as the one-week-old Sol, and the cache-read discount moved from 90 percent to 95 percent. Artificial Analysis translates those rates into a cost-to-run figure by holding the Intelligence Index constant and reporting the dollar bill for the same task on each model: $0.72 for GPT-6.1 Sol against $3.26 for GPT-6 Astra. On output tokens the newer Sol uses between 10 and 30 percent more than the model it replaced, so part of what a router gains on index points is spent back on a longer answer at the same per-token price.

For a team already paying Astra prices on Codex, Copilot or ChatGPT Work tickets, the second source removes the main objection to a routing change: OpenAI's own 1.9 percent factual-error gap at launch was a single vendor number, and the Artificial Analysis run on a different test family lands close to it. The move the report suggests, read literally, is to run Sol as the default and keep Astra on the tasks where one Intelligence Index point actually changes the answer, which a team can only name from its own tickets.

Source

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX