Skip to content
OddBrief
AI2 minPrimary source linked

GitHub's ReviewBench is graded by Claude and only shows a tool's best score

GitHub released an open benchmark for AI code review built from 219 pull requests. Its judge is an Anthropic model, and its leaderboard only updates upward.

AI-assisted, human-reviewed

Illustration of printed code pages with handwritten review marks under a desk lampAI
AI-generated illustration

Key facts

Who
GitHub and Microsoft researchers
What
ReviewBench, an open benchmark of 219 pull requests from 187 repositories in 19 languages
Judge
Claude Sonnet 5 grades all submissions; leaderboard posts a score only if it is an improvement
Name clash
LangChain published a different ReviewBench on July 31, 2026

GitHub has released ReviewBench, a public test for AI code reviewers built from 219 public pull requests, and says it already uses it to predict how changes to Copilot code review will perform. Two choices in its design stand out: the grader is Anthropic's Claude Sonnet 5, and the public leaderboard only records a tool's new score if it beats the old one.

The benchmark, announced on October 5 and available as a research preview, also arrives two months after LangChain published its own code review benchmark under the same name.

How GitHub built it

GitHub says it analyzed 103.9 million pull requests to match real-world workloads. The final set covers 219 pull requests from 187 open source repositories in 19 languages, with language and repository size mirroring GitHub overall. Pull request size is deliberately tilted away from tiny one-file changes toward larger ones where review matters.

The list of known issues for each pull request was assembled from human reviewers' comments, fixes authors made later, static analysis tools and several frontier models, then deduplicated and judged against one rubric. Senior engineers who had not built the dataset re-labelled every finding from scratch and agreed with it 96.6% of the time, according to GitHub.

A judge from another company, and an upward-only board

GitHub names Claude Sonnet 5 as the LLM grader and publishes the rubric and judge configuration. That means the scores of every tool, including Copilot's competitors and tools built on Claude, depend on one model's view of what counts as a real issue.

Submitters run their tool on all 219 pull requests three times. Scores stay private until a maintainer approves them, and they are published only if they beat the agent's current leaderboard score or are its first entry. A tool that gets worse in a later version will not show that on the board.

More comments, slightly better ones

GitHub's example of the benchmark working is telling. An ensemble that merged several model runs into one review raised the share of comments developers acted on by 8.0% and recall by 13.6% in a live A/B test, while comment volume rose 61% and cost per review fell 8.0%. Critical comments rose 262%, close to the 227% the benchmark had predicted.

LangChain's version, published July 31, took a different route: 59 tasks built from comments by trusted reviewers in its own LangSmith repository. With a basic setup, the strongest models recovered only about 30% of those issues, LangChain said.

GitHub invites others to challenge its assumptions. Whether rival vendors submit to a benchmark run and judged on terms set by Copilot's maker is the next test.

Sources