GitHub ReviewBench compares AI code review agents on 219 public pull requests from 187 repositories across 19 languages. Launched on 2026-10-05, it combines a human-reviewed golden set with a Claude Sonnet 5 classifier and reports grounded and augmented precision, recall, and F1 scores.
ReviewBench key facts
- GitHub modeled the corpus on 103.9 million GitHub pull requests.
- The corpus contains 219 pull requests from 187 public open source licensed repositories across 19 languages.
- The published methodology names Claude Sonnet 5 as the classifier. It says the corpus was initially labeled with Claude Sonnet 4.6 and later updated.
- A senior-engineer audit agreed with the classifier's true-positive and false-positive labels 96.6% of the time. It agreed within one severity level 98.7% of the time, exactly on severity 62.9% of the time, and exactly on category 79.7% of the time.
- The audit manually corrected 47 findings that the classifier had labeled as true positives.
- Grounded recall is the headline cross-agent comparison. Augmented recall is a per-agent diagnostic because each agent's newly discovered true positives change its denominator.
- GitHub's 2026-10-05 launch post was written by Michelle Zhou and Alejandro Carderera de Diego.
ReviewBench corpus design
GitHub matched the corpus's language and repository-size distributions to its wider pull-request population. It deliberately weighted pull-request size toward the reviewable middle and tail, so large and multi-file changes appear more often than they would in an unmodified sample.
Added and removed lines | Pull requests | Share |
|---|---|---|
50 or fewer | 17 | 7.8% |
51–200 | 40 | 18.3% |
201–500 | 44 | 20.1% |
501–1,000 | 40 | 18.3% |
More than 1,000 | 78 | 35.6% |
Total | 219 | 100% |
The primary languages are TypeScript (68 pull requests), Python (41), C# (25), Go (19), and JavaScript (15). The other 14 languages account for 51 pull requests. The largest repository contributes 10 pull requests, or 4.6% of the corpus. ReviewBench infers pull-request types from titles and descriptions rather than treating them as formal human labels: 79 are features, 59 bug fixes, 15 documentation changes, 14 performance changes, 12 refactors, and 40 other changes.
ReviewBench ground truth and scoring
ReviewBench builds its golden set in three stages:
- It collects candidate findings from human reviewers, issues inferred from author follow-up commits, deterministic analysis tools, and multiple frontier LLMs.
- It semantically deduplicates findings that describe the same underlying issue.
- It applies a shared rubric with the Claude Sonnet 5 classifier.
The methodology defines a true positive as a finding that is factually correct and verifiable against the changed code, relevant to improving the pull request, and within the scope of the review. The classifier also assigns severity, category, scope, difficulty, context required, and actionability labels.
Grounded ReviewBench metrics
Grounded precision, recall, and F1 use the fixed human-reviewed golden set. They provide the apples-to-apples comparison because every agent is measured against the same known issues.
Augmented ReviewBench metrics
Augmented precision, recall, and F1 also classify findings that do not match the golden set. A valid new finding can earn credit, but each agent's new true positives expand its recall denominator. ReviewBench therefore treats augmented recall as a per-agent diagnostic and grounded recall as the cross-agent headline.
ReviewBench leaderboard on 2026-10-06
The ReviewBench site says its initial entries were produced by the ReviewBench team and published on 2026-10-05. Vendors did not verify those entries, and the products may have changed since their listed run dates. These are the leaderboard values shown on 2026-10-06.
Rank | Agent | Grounded F1 | Grounded precision | Grounded recall | Augmented F1 |
|---|---|---|---|---|---|
1 | Copilot Code Review Balanced | 40.1% | 87.8% ± 0.5 | 26.0% ± 0.6 | 49.7% |
2 | Devin AI | 37.0% | 84.0% ± 0.7 | 23.8% ± 0.1 | 47.3% |
3 | Qodo | 35.1% | 85.3% ± 0.8 | 22.1% ± 0.9 | 44.5% |
4 | Codex GPT-5.6 Sol · Ultra | 32.3% | 87.0% ± 2.0 | 19.9% ± 0.1 | 43.2% |
5 | Cubic | 27.3% | 85.5% | 16.3% | 37.4% |
6 | Greptile | 27.2% | 86.1% | 16.2% | 32.8% |
7 | Cursor | 17.3% | 87.7% ± 0.8 | 9.6% ± 0.3 | 21.7% |
ReviewBench and the Copilot code review A/B test
GitHub defines addressed rate as the percentage of Copilot code review comments that an LLM determines prompted a developer to make a corresponding code change. Its online recall signal measures how much additional human review is still needed. For a lite-tier multi-model ensemble, the online A/B test moved in the direction ReviewBench predicted.
Metric | ReviewBench offline prediction | Online result in the GitHub post |
|---|---|---|
Precision, measured online as addressed rate | +4.45% | +8.00% |
Recall | +12.88% | +13.58% |
Comment volume | +40.0% | +61% in the post's prose; +25.0% in its chart text |
Cost per review | -23.7% | -8.00% |
The post also reports critical comments at +227.4% offline and +262.0% online, moderate comments at +121.0% and +18.00%, and low-severity comments at -15.69% and -45.0%. The 61% and 25.0% online comment-volume values are both present in the same post, so they should not be presented as one uncontested number.
Running ReviewBench locally and submitting a reviewer
ReviewBench's repository provides a local container check and a self-service submission flow:
~~~bash
git clone https://github.com/review-bench/ReviewBench && cd ReviewBench
scripts/try-agent.sh my-reviewer:dev -e MYAPIKEY
scripts/try-agent.sh my-reviewer:dev --set full -e MYAPIKEY
~~~
The first command runs the 25-pull-request test set. The second runs all 219 pull requests. These local checks validate that the container produces a valid findings file; they do not create a leaderboard score. Official submission requires registering the container image and configuration at https://review-bench.ai/submit. The final evaluation runs three rounds over all 219 pull requests with the benchmark judge, and a maintainer approves the result before publication.
ReviewBench sources
- ReviewBench: An open benchmark for AI code review: https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/ (read 2026-10-06)
- ReviewBench leaderboard: https://review-bench.ai/ (read 2026-10-06)
- ReviewBench methodology: https://review-bench.ai/methodology (read 2026-10-06)
- ReviewBench GitHub repository: https://github.com/review-bench/ReviewBench (read 2026-10-06)
Last verified: 2026-10-06.