AI NewsDev toolsReported

An independent code-review benchmark ranks Copilot fourth, behind Cubic

The New Stack reports that Martian's Code Review Bench ranks Cubic first and Copilot fourth, while GitHub's own ReviewBench, released this week, puts Copilot on top.

AI News

Editorial2 min read

LinkedInX

Why it mattersA team picking an AI code reviewer should check at least one benchmark the vendor does not run itself, because the ranking changes when somebody else measures.

The code reviewer that ranks first on a benchmark its own vendor publishes can rank fourth on a benchmark somebody else runs.

The New Stack's Paul Sawers reports that Martian's Code Review Bench, run by an AI research lab that does not sell coding tools, puts Cubic first and GitHub Copilot fourth on its online leaderboard as of October 6. GitHub published its own benchmark, ReviewBench, on the same site this week, and that leaderboard puts Copilot first.

What GitHub measures

ReviewBench uses 219 pull requests from 187 public repositories across 19 programming languages, selected from a set of 103.9 million pull requests. A reviewer scores on a grounded F1, which combines precision and recall against a reference set built from human review comments, author follow-up changes, static-analysis tools and language-model reviewers. The New Stack notes that Claude Sonnet 5 classifies the findings, and that human and classifier judgments agreed 96.6 percent of the time on whether a finding was a true or false positive. Copilot in its Balanced configuration leads the leaderboard with a 40.1 percent grounded F1.

Two caveats apply to the GitHub result, both stated in the piece. GitHub generated every initial entry itself by running the publicly available version of each product, and the vendors did not run or verify those tests. The products were also tested on different dates: Copilot on October 1, Cubic and Greptile back in June, so the Copilot result is scored against a more recent build than the others.

What Martian measures

Martian's online leaderboard uses a different signal. The New Stack writes that it watches how developers respond to review comments across real open-source repositories, so a tool's score reflects whether its suggestions led to code changes rather than whether its comments matched a reference set. On that leaderboard, Cubic sits first with a 64.9 percent F1, Greptile second, CodeRabbit third, and Copilot fourth at 60.9. Martian also publishes an offline leaderboard, where Qodo Deep ranks first, followed by Cubic and Augment, and Copilot sits fifth with a 58 percent F2 score. F2 weights recall more heavily than precision.

Martian's benchmark and methodology are released under an MIT licence through its public repository, and its February launch post framed its own role as a third option alongside academic and vendor-produced benchmarks: a lab that does not train models or sell coding tools and therefore has no stake in which tool wins.

The two scores are not directly comparable. The datasets are different, the scoring methods are different, and Martian's online and offline tests measure different kinds of evidence. The useful fact for a reader is this: the ranking of AI code reviewers changes when the person running the benchmark changes. A buying decision based on one leaderboard, especially one published by a vendor whose own product appears on it, is based on less evidence than the number suggests.

Source

The New Stack, Copilot tops GitHub's own AI code review benchmark. An independent one tells a different story. by Paul Sawers. Martian's live Code Review Bench leaderboard. GitHub's ReviewBench methodology.

Reported byThe New Stack

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX