CodeRabbit put GPT-6 Astra through code review and found 2.3 points of coverage for 2.5 times the price

Image: CodeRabbit
Why it mattersThe newest model wins on coverage by about two points and costs two and a half times as much, so the choice is a budget decision per repository rather than a default upgrade.
CodeRabbit, which sells an automated code review product, published an evaluation on 4 September comparing OpenAI's GPT-6 Astra against GPT-5.6 Sol and Opus 5 on its own review workload. The company reports Astra found the most bugs and costs the most to run, and the gap in each direction is not the same size.
The numbers CodeRabbit reports
On overall actionable bug coverage, CodeRabbit puts Astra at 61.3%, GPT-5.6 Sol at 59.0% and Opus 5 at 50.2%. On cross-file review, where a defect only shows up if the model connects changes in more than one file, it reports Astra at 57.1%, Sol at 47.6% and Opus 5 at 42.9%.
Expressed as relative gains, CodeRabbit says Astra caught approximately 4% more labelled bugs through actionable findings than Sol and 22% more than Opus 5, rising to 20% over Sol and 33% over Opus 5 on the harder cross-file reviews.
The published prices, per million tokens, put Astra at $10.00 input and $50.00 output. Sol is $4.00 and $20.00. GPT-5.6 Terra is $2.00 and $12.00, GPT-5.6 Luna is $0.20 and $1.20, and Claude Fable 5.1 matches Astra at $10.00 and $50.00. On CodeRabbit's illustrative task of 100,000 input and 10,000 output tokens, Astra costs $1.50 against Sol's $0.60.
What the evaluation does not say
CodeRabbit does not publish the size or composition of the dataset, and does not give latency figures or false positive rates. It calls the result "an early, directional result" and says the numbers "do not establish an overall ranking of review quality, predict a team's defect rate, or promise the same gain on every pull request."
Every figure above is CodeRabbit's, measured on CodeRabbit's own harness. The company does not sell any of the models it compared, so it has no stake in which one wins, but a benchmark with no published dataset cannot be reproduced by anyone else.
CodeRabbit also states that neither it nor its model providers train on customer code, and that Astra supports zero data retention for eligible API customers.
Two points of coverage for two and a half times the price is a real trade, and which way it falls depends on the repository rather than on the model. A codebase where the expensive misses are cross-file, where a change in one module quietly breaks a caller in another, is where Astra's largest reported margin sits, at 57.1% against Sol's 47.6%. A codebase where most review findings are local gets far less for the extra spend.
The other reading is the one the price table makes hard to ignore. Luna is fifty times cheaper on input than Astra, and CodeRabbit did not publish a coverage number for it. Anyone deciding what to point at a review queue wants that row filled in, because the question is rarely whether the best model is best. It is how much coverage the cheap one already buys.
Source
- GPT-6 Astra in code review: Gains, privacy, and cost, CodeRabbit, 4 September 2026
Source: CodeRabbit
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually shipping with them. Short, and only when there is something worth reading.
