
Why it mattersAn open score lets a team compare AI code reviewers on the same pull requests, so a buying decision can be made from numbers rather than a vendor demo.
Picking an AI code reviewer today means watching a demo and guessing.
GitHub released ReviewBench on 5 October 2026, an open code review benchmark drawn from 219 pull requests across 187 open-source licensed repositories, in 19 programming languages. The authors, Michelle Zhou and Alejandro Carderera de Diego, say the pull requests were selected from a dataset of 103.9 million public pull requests, with multi-source ground truth and a calibrated evaluation.
What it measures
ReviewBench scores a reviewer on six numbers: grounded precision, grounded recall and a grounded F1, plus the same three in an augmented form. Each finding carries a severity tag of Critical, Medium or Low, and a category tag drawn from Correctness, Security, Reliability, Maintainability and Testing. The grounded set checks whether the reviewer catches the thing a human reviewer caught in the same pull request; the augmented set lets a reviewer earn credit for a real issue the human missed.
The benchmark lives at review-bench.ai and GitHub says the dataset is publicly available.
What the result is used for
GitHub says it tuned its own Copilot Code Review against ReviewBench. The post reports a single example, an iteration that lifted the addressed rate by 8.0 percent and recall by 13.6 percent, which are GitHub's own figures. Other reviewers were not scored in the post, so there is no cross-vendor comparison today.
That is the point of releasing the benchmark. A team deciding between a Copilot review, an in-house agent and a third-party reviewer can run the same 219 pull requests through each one and read the six numbers side by side, instead of judging on a demo of two or three handpicked diffs.
Severity tags matter here as much as the raw score. A reviewer that catches many Low findings but misses the one Critical issue is worse than the opposite; a flat F1 across all severities would hide that. The ground-truth labels mean a buyer can filter down to Critical findings in the Security and Correctness categories and compare only on those.
Two limits are worth stating. The pull requests are open source and licensed, which means the benchmark skews toward public code with public review history, and a reviewer may behave differently on a private codebase it has never seen. And the 103.9 million number is the pool the 219 were drawn from, not the test set itself; the test set is small enough that a single-run score has sampling noise, and a buyer should run it more than once.
Source
- ReviewBench: an open benchmark for AI code review, GitHub, 5 October 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.


