GitHub ReviewBench scores AI code review agents on 219 pull requests

GitHub ReviewBench compares AI code review agents on 219 pull requests with grounded and augmented precision, recall, and F1 metrics.

GitHub ReviewBench compares AI code review agents on 219 public pull requests from 187 repositories across 19 languages. Launched on 2026-10-05, it combines a human-reviewed golden set with a Claude Sonnet 5 classifier and reports grounded and augmented precision, recall, and F1 scores.

ReviewBench key facts

  • GitHub modeled the corpus on 103.9 million GitHub pull requests.
  • The corpus contains 219 pull requests from 187 public open source licensed repositories across 19 languages.
  • The published methodology names Claude Sonnet 5 as the classifier. It says the corpus was initially labeled with Claude Sonnet 4.6 and later updated.
  • A senior-engineer audit agreed with the classifier's true-positive and false-positive labels 96.6% of the time. It agreed within one severity level 98.7% of the time, exactly on severity 62.9% of the time, and exactly on category 79.7% of the time.
  • The audit manually corrected 47 findings that the classifier had labeled as true positives.
  • Grounded recall is the headline cross-agent comparison. Augmented recall is a per-agent diagnostic because each agent's newly discovered true positives change its denominator.
  • GitHub's 2026-10-05 launch post was written by Michelle Zhou and Alejandro Carderera de Diego.

ReviewBench corpus design

GitHub matched the corpus's language and repository-size distributions to its wider pull-request population. It deliberately weighted pull-request size toward the reviewable middle and tail, so large and multi-file changes appear more often than they would in an unmodified sample.

Added and removed lines

Pull requests

Share

50 or fewer

17

7.8%

51–200

40

18.3%

201–500

44

20.1%

501–1,000

40

18.3%

More than 1,000

78

35.6%

Total

219

100%

The primary languages are TypeScript (68 pull requests), Python (41), C# (25), Go (19), and JavaScript (15). The other 14 languages account for 51 pull requests. The largest repository contributes 10 pull requests, or 4.6% of the corpus. ReviewBench infers pull-request types from titles and descriptions rather than treating them as formal human labels: 79 are features, 59 bug fixes, 15 documentation changes, 14 performance changes, 12 refactors, and 40 other changes.

ReviewBench ground truth and scoring

ReviewBench builds its golden set in three stages:

  1. It collects candidate findings from human reviewers, issues inferred from author follow-up commits, deterministic analysis tools, and multiple frontier LLMs.
  2. It semantically deduplicates findings that describe the same underlying issue.
  3. It applies a shared rubric with the Claude Sonnet 5 classifier.

The methodology defines a true positive as a finding that is factually correct and verifiable against the changed code, relevant to improving the pull request, and within the scope of the review. The classifier also assigns severity, category, scope, difficulty, context required, and actionability labels.

Grounded ReviewBench metrics

Grounded precision, recall, and F1 use the fixed human-reviewed golden set. They provide the apples-to-apples comparison because every agent is measured against the same known issues.

Augmented ReviewBench metrics

Augmented precision, recall, and F1 also classify findings that do not match the golden set. A valid new finding can earn credit, but each agent's new true positives expand its recall denominator. ReviewBench therefore treats augmented recall as a per-agent diagnostic and grounded recall as the cross-agent headline.

ReviewBench leaderboard on 2026-10-06

The ReviewBench site says its initial entries were produced by the ReviewBench team and published on 2026-10-05. Vendors did not verify those entries, and the products may have changed since their listed run dates. These are the leaderboard values shown on 2026-10-06.

Rank

Agent

Grounded F1

Grounded precision

Grounded recall

Augmented F1

1

Copilot Code Review Balanced

40.1%

87.8% ± 0.5

26.0% ± 0.6

49.7%

2

Devin AI

37.0%

84.0% ± 0.7

23.8% ± 0.1

47.3%

3

Qodo

35.1%

85.3% ± 0.8

22.1% ± 0.9

44.5%

4

Codex GPT-5.6 Sol · Ultra

32.3%

87.0% ± 2.0

19.9% ± 0.1

43.2%

5

Cubic

27.3%

85.5%

16.3%

37.4%

6

Greptile

27.2%

86.1%

16.2%

32.8%

7

Cursor

17.3%

87.7% ± 0.8

9.6% ± 0.3

21.7%

ReviewBench and the Copilot code review A/B test

GitHub defines addressed rate as the percentage of Copilot code review comments that an LLM determines prompted a developer to make a corresponding code change. Its online recall signal measures how much additional human review is still needed. For a lite-tier multi-model ensemble, the online A/B test moved in the direction ReviewBench predicted.

Metric

ReviewBench offline prediction

Online result in the GitHub post

Precision, measured online as addressed rate

+4.45%

+8.00%

Recall

+12.88%

+13.58%

Comment volume

+40.0%

+61% in the post's prose; +25.0% in its chart text

Cost per review

-23.7%

-8.00%

The post also reports critical comments at +227.4% offline and +262.0% online, moderate comments at +121.0% and +18.00%, and low-severity comments at -15.69% and -45.0%. The 61% and 25.0% online comment-volume values are both present in the same post, so they should not be presented as one uncontested number.

Running ReviewBench locally and submitting a reviewer

ReviewBench's repository provides a local container check and a self-service submission flow:

~~~bash

git clone https://github.com/review-bench/ReviewBench && cd ReviewBench

scripts/try-agent.sh my-reviewer:dev -e MYAPIKEY

scripts/try-agent.sh my-reviewer:dev --set full -e MYAPIKEY

~~~

The first command runs the 25-pull-request test set. The second runs all 219 pull requests. These local checks validate that the container produces a valid findings file; they do not create a leaderboard score. Official submission requires registering the container image and configuration at https://review-bench.ai/submit. The final evaluation runs three rounds over all 219 pull requests with the benchmark judge, and a maintainer approves the result before publication.

ReviewBench sources

Last verified: 2026-10-06.

Spotted an outdated or wrong claim? Agents can report it with evidence throughPOST /api/feedback; an editor checks every report. See llms.txt for the agent API.