AI NewsModels & agentsReported

Artificial Analysis says Gemini 4 Argon ties GPT-6 Astra at 53 on its Intelligence Index and costs 39 percent less per task

Artificial Analysis published independent benchmark numbers for Google's new Gemini 4 Argon (High) model, showing it scored 53 on the firm's Intelligence Index, the same as GPT-6 Astra (Max), while costing $1.99 per Intelligence Index task against Astra's $3.26.

AI News

Editorial2 min read

LinkedInX

Why it mattersA team picking a top-tier model for an agent run that uses many tokens now has an independent cost-per-task comparison with the same intelligence rating on both sides, which is the number that decides which model goes into production for a long-running job.

A team routing a long agent task to a top-tier model decides on the pair of numbers the vendor rarely puts side by side: how well the model scores on a shared benchmark, and what one benchmark run actually costs. Artificial Analysis published both for Google's new Gemini 4 Argon on 30 September, nine hours after the Google launch, and the independent numbers put Argon in a precise position against OpenAI's current flagship.

On Artificial Analysis's own Intelligence Index v4.3.2, which combines ten evaluations including Humanity's Last Exam, GDPval-AA, Terminal-Bench 4.0, SciCode and AA-Omniscience, Argon (High) scored 53. The median across the models Artificial Analysis tracks is 26. GPT-6 Astra (Max), OpenAI's current flagship, also scored 53. So on intelligence the two tie.

The cost per task is where the comparison opens

Artificial Analysis reports Argon (High) at $2.00 per 1M input tokens and $10.00 per 1M output tokens, which matches the median price for its intelligence tier. Astra (Max) is $10.00 input and $50.00 output. Running the Intelligence Index itself cost Artificial Analysis $1.99 for Argon and $3.26 for Astra, so Argon came in 39 percent cheaper per task at the same intelligence score. The firm notes Argon has a 1M-token context window and accepts text and image input.

Argon's lower cost comes partly from the lower price per token and partly from how many tokens it generated. Artificial Analysis measured Argon producing 110M tokens running the Index, against Astra's 60M. The median across tested models is 82M. So Argon is wordier for the same answer quality: Artificial Analysis labels the output 28M tokens above median. A team paying by the token on long-running agent jobs should measure this on its own prompts before switching, because the per-task cost on their workload may be closer to Astra's than to Argon's.

What the number does not say

Artificial Analysis's methodology note stresses the tests are run on dedicated hardware with a weighted average of ten evaluations, so the Index is one number from many. The firm does not report side-by-side benchmarks on tool-use latency, agent trajectory stability or safety refusal rate, each of which can change which model a team actually ships. The 53-vs-53 result also tells a team nothing about which model is better on their own code or their own documents, and Google's launch material (which named cybersecurity and coding as Argon's target work) does not confirm the Intelligence Index score at all.

For a team already running Astra as its default model, these numbers justify an A/B test on real traffic. The intelligence rating is the same, the per-task cost is lower, and the comparison is between the vendor-labelled thinking variants of each flagship. For a team on a cheaper model that pays per-token on long agent runs, the 110M-token figure is the one to measure first on your own workload, because the cost advantage on paper disappears if your prompts push Argon into producing more output than the Index did.

Source

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX