GitHub ReviewBench: AI Code Review Leaderboard, Scores and How to Test
Quick answer: GitHub’s new ReviewBench is an open research-preview benchmark for AI code review agents. It uses 219 public pull requests from 187 open-source repositories across 19 programming languages, selected from distributions modeled after 103.9 million GitHub pull requests. The benchmark measures how accurately code-review agents identify real issues while controlling noise, and its current leaderboard puts GitHub Copilot Code Review’s Balanced configuration at the top on grounded recall.
ReviewBench matters because code-review agents can look impressive in demos while behaving very differently on real pull requests. One may catch more bugs but produce too much noise; another may be precise but miss important findings. GitHub designed the benchmark to make those tradeoffs visible with a reproducible dataset, shared rubric and common scoring pipeline.
ReviewBench at a glance
| Item | Current ReviewBench design |
|---|---|
| Corpus | 219 public pull requests |
| Repositories | 187 open-source repositories |
| Languages | 19 |
| Reference distribution | 103.9 million GitHub pull requests |
| Ground truth | Human reviews, follow-up fixes, static analysis and frontier LLM findings |
| Judge | Claude Sonnet 5 under a published rubric |
| Senior-engineer agreement | 96.6% on independently relabeled golden true positives |
| Main comparison metric | Grounded recall, alongside precision and F1 |
Current ReviewBench leaderboard
The initial public results show a striking pattern: the leading tools are all fairly precise, but they differ much more on recall. In other words, the main gap is not whether a reported issue is real; it is how many known issues each reviewer manages to find.
| Reviewer | Grounded precision | Grounded recall | Grounded F1 |
|---|---|---|---|
| Copilot Code Review (Balanced) | 87.8% | 26.0% | 40.1% |
| Devin | 84.0% | 23.8% | 37.0% |
| Qodo | 85.3% | 22.1% | 35.1% |
| Codex (GPT-5.6 Sol · Ultra) | 87.0% | 19.9% | 32.3% |
| Cubic | 85.5% | 16.3% | 27.3% |
| Greptile | 86.1% | 16.2% | 27.2% |
| Cursor | 87.7% | 9.6% | 17.3% |
These are benchmark results, not a universal ranking of developer tools. GitHub built ReviewBench and its own Copilot configuration currently leads the table, so independent reproduction matters. GitHub also notes that the initial entries were run by the ReviewBench team rather than the vendors themselves.
Why recall is the headline metric
If a reviewer produces ten comments and nine are valid, its precision is high. But if the pull request contains 30 known issues and the reviewer finds only three, recall is low. A developer can experience the reviewer as “quiet and accurate” while still missing most problems.
ReviewBench therefore reports both precision and recall rather than collapsing quality into a single subjective score. It also lets users adjust the beta term in F-beta scoring when they care more about broad coverage or more about minimizing noise.
How GitHub built the benchmark corpus
GitHub analyzed 103.9 million pull requests to understand real-world distributions by language, repository size and change shape. ReviewBench then samples 219 public pull requests from 187 open-source licensed repositories across 19 languages.
The benchmark intentionally does not mirror tiny one-file changes perfectly. GitHub weights the corpus toward reviewable middle and larger changes so the benchmark contains substantive cases where code review quality matters.
How the golden set is created
A benchmark is only as trustworthy as its reference findings. GitHub avoids relying on a single source by gathering candidate issues from several places:
- real human review comments;
- issues inferred from follow-up fix commits made by the authors;
- deterministic static-analysis tools;
- multiple frontier language models from different model families.
Overlapping findings are semantically deduplicated, then judged against one rubric. A finding counts only if it is true, relevant and non-trivial. GitHub uses Claude Sonnet 5 as the automated judge and publishes the evaluation rubric and judge configuration.
Why the 96.6% agreement number matters
Senior engineers who were not involved in building the dataset independently relabeled the golden true positives. Their judgments agreed with ReviewBench 96.6% of the time. That does not make the benchmark infallible, but it gives users a concrete signal about the quality of the reference labels.
Grounded vs augmented metrics
ReviewBench has an important feature that many fixed benchmarks lack. A reviewer can discover a real issue that was not already in the golden set. If the benchmark automatically treated every unmatched finding as false, stronger systems would eventually be punished for finding new bugs.
To address that, ReviewBench separates two metric families:
- Grounded precision, recall and F1: compare only against the known golden-set labels.
- Augmented precision, recall and F1: allow the judge to evaluate unmatched findings and give credit when they are genuinely valid.
GitHub uses grounded recall as the main cross-system comparison because the denominator stays fixed across reviewers. Augmented metrics remain useful diagnostics for systems that uncover issues the original producers missed.
Severity and category breakdowns
A single score can hide what a reviewer is actually good at. ReviewBench labels findings by severity and category, allowing teams to compare tools on the types of errors they care about.
Published categories include correctness, security, reliability, maintainability and testing, while severity distinguishes critical, medium and low-priority issues. A security-focused team may prefer a reviewer that catches more serious vulnerabilities even if its overall F1 is not the highest.
Does ReviewBench predict production quality?
GitHub says its offline benchmark has tracked the direction of later production experiments. In one Lite-tier experiment, a multi-model ensemble increased production recall by 13.6% and lowered cost per review by 8.0%, moving in the direction ReviewBench predicted.
GitHub also reported an 8.0% increase in its online addressed-rate proxy for precision and a 61% increase in comment volume. The company is careful to note that online experiments remain the ultimate measure of user impact; the offline benchmark is a faster signal for deciding which changes are worth testing.
How to submit your own AI code reviewer
- Sign in to the ReviewBench site with GitHub.
- Register the reviewer, including a container image, configuration and your model key.
- Run the 25-pull-request test set while tuning your agent.
- Run the full set of 219 pull requests for three rounds.
- Submit the score for maintainer review before publication to the leaderboard.
The score stays private until a maintainer approves it. The benchmark publishes a new result when it is the agent’s first entry or improves its existing leaderboard score.
What ReviewBench does well
Its strongest feature is auditability. The dataset, rubric, judge configuration and runner are available so researchers and vendors can inspect the evaluation rather than treating a leaderboard as a black box. The use of representative pull requests also makes the test more relevant than toy bug snippets.
What ReviewBench cannot tell you
No benchmark can tell you which reviewer is automatically best for every repository. Results can shift with language mix, codebase architecture, organization-specific conventions, security requirements, model versions and configuration.
The benchmark also measures findings, not the complete developer experience. Latency, pricing, repository indexing, comment presentation, IDE integration, privacy controls and how easily a developer acts on a suggestion still matter in production.
How teams should use the leaderboard
Use ReviewBench as a screening signal, then validate the finalists on a private set of your own historical pull requests. Measure both how many meaningful issues are found and how many low-value comments engineers dismiss. A tool with slightly lower benchmark recall may still be better for a specific organization if it understands the repository and fits the team’s review process.
For related developer-agent changes, read AVARIXO’s guide to GPT-6 Astra and GPT-6.1 Sol performance and the llama.cpp v0.6.0 update.
FAQ
Is ReviewBench open?
GitHub says the complete dataset, evaluation methodology, judge prompt, judge configuration and self-serve runner are available for research and submissions.
How many pull requests are in ReviewBench?
The current research-preview corpus contains 219 public pull requests from 187 repositories across 19 languages.
Which reviewer currently has the highest grounded recall?
In the initial leaderboard reported by GitHub, Copilot Code Review Balanced leads grounded recall at 26.0%.
Why is grounded recall only 26% at the top?
The benchmark is intentionally demanding and uses a broad multi-source golden set. High precision combined with lower recall indicates that current reviewers generally avoid many false alarms but still miss a large share of known findings.
Can I test my own review agent?
Yes. ReviewBench provides an onboarding path, a 25-PR test set and a full 219-PR evaluation.
Sources
Primary source: GitHub — ReviewBench: An open benchmark for AI code review, published October 5, 2026. Current leaderboard figures were cross-checked against the research-preview reporting on October 7, 2026.
