<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>ai-review-agents on tomrochette.com</title>
    <link>https://tomrochette.com/tags/ai-review-agents/</link>
    <description>Recent content in ai-review-agents on tomrochette.com</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <managingEditor>tom@tomrochette.com (Tom Rochette)</managingEditor>
    <webMaster>tom@tomrochette.com (Tom Rochette)</webMaster>
    <copyright>© 2026 Tom Rochette</copyright>
    <lastBuildDate>Wed, 07 Oct 2026 06:38:49 -0400</lastBuildDate><atom:link href="https://tomrochette.com/tags/ai-review-agents/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>ReviewBench</title>
      <link>https://tomrochette.com/agents/evaluation-review/reviewbench/</link>
      <pubDate>Wed, 07 Oct 2026 00:00:00 +0000</pubDate>
      <author>tom@tomrochette.com (Tom Rochette)</author>
      <guid>https://tomrochette.com/agents/evaluation-review/reviewbench/</guid>
      <category>research-note</category><category>agent-curated</category><category>fully-ai-generated</category><category>llm=glm-5.3-flash</category><category>benchmark</category><category>evaluation</category><category>code-review</category><category>ai-review-agents</category>
      <description>&lt;p&gt;ReviewBench is GitHub&amp;rsquo;s open offline benchmark for AI code review agents: 219 pull requests sampled from the public GitHub corpus, a multi-source golden set, a published rubric graded by Claude Sonnet 5, and a public leaderboard at review-bench.ai.&lt;/p&gt;&#xA;&lt;p&gt;&lt;strong&gt;ReviewBench is the first code-review benchmark built by the company that sells the leading reviewer, so its value rests entirely on the parts it publishes (the dataset, the rubric, the judge, the self-serve runner) and the part it keeps, a maintainer approval gate before any score goes public, is where a reader has to watch hardest.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;What it is&#xA;    &lt;div id=&#34;what-it-is&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#what-it-is&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A research preview built by GitHub with Microsoft and announced on 2026-10-05, with the corpus, methodology, judge prompt, and runner in an MIT-licensed repository.&#xA;The benchmark sampled 219 public pull requests from 187 open source licensed repositories across 19 languages, choosing them to match the language and repository-size distribution of 103.9 million GitHub pull requests while weighting PR size toward the reviewable middle.&#xA;Ground truth comes from human review comments, issues inferred from author follow-up commits, deterministic analysis tools, and multiple frontier LLMs, deduplicated and validated under one rubric, with Claude Sonnet 5 as the grader and a separate matcher deciding whether a candidate finding corresponds to a golden one.&#xA;Six metrics score each system: grounded precision, recall, and F1 against the golden set, and augmented versions that also judge unmatched findings, with grounded recall as the headline cross-system comparison and an adjustable Fβ weight for precision-versus-recall preferences.&#xA;Submission is self-serve: sign in with GitHub, register a container image and your own model key, iterate on a 25-PR test set, then run the full 219 three times under the same judge as every other entry.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Status&#xA;    &lt;div id=&#34;status&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#status&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;A four-day-old research preview: the repository (review-bench/ReviewBench) was created 2026-09-04 and sits at 26 stars and 5 open issues as of 2026-10-07, last pushed 2026-10-06.&lt;/p&gt;&#xA;&lt;picture&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: dark)&#34; srcset=&#34;https://api.star-history.com/chart?repos=review-bench/ReviewBench&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;source media=&#34;(prefers-color-scheme: light)&#34; srcset=&#34;https://api.star-history.com/chart?repos=review-bench/ReviewBench&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;  &lt;img alt=&#34;Star History Chart&#34; src=&#34;https://api.star-history.com/chart?repos=review-bench/ReviewBench&amp;type=date&amp;theme=dark&amp;legend=top-left&#34; /&gt;&#xA;&lt;/picture&gt;&#xA;&lt;p&gt;Validation is the strongest published number: senior engineers who did not build the dataset re-labeled every ground-truth finding and agreed with it 96.6 percent of the time, with 47 findings manually corrected, and GitHub reports the benchmark&amp;rsquo;s offline movement has anticipated production A/B directions for Copilot code review (the lite-tier ensemble experiment: addressed rate up 8.0 percent, recall up 13.6 percent, cost per review down 8.0 percent).&#xA;&lt;strong&gt;The inaugural leaderboard&amp;rsquo;s top entry is GitHub&amp;rsquo;s own Copilot code review at 40.1 grounded F1 in its Balanced configuration as of 2026-10-06, and the Hacker News footprint is a 4-point, zero-comment thread from launch day, so the benchmark&amp;rsquo;s reach so far is press and vendor channels, not community debate.&lt;/strong&gt;&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Strengths&#xA;    &lt;div id=&#34;strengths&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#strengths&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Everything needed to reproduce or contest a score is public: the corpus and golden findings, repository mirrors, the rubric, the judge prompt and model configuration, the matcher, and the agent contract.&lt;/li&gt;&#xA;&lt;li&gt;The golden set does not depend on any one producer&amp;rsquo;s blind spots, and the augmented metrics give credit for valid findings its creators did not anticipate.&lt;/li&gt;&#xA;&lt;li&gt;The Fβ knob plus severity and category slicing let a team re-rank the leaderboard by its own review preference instead of accepting one number.&lt;/li&gt;&#xA;&lt;li&gt;GitHub&amp;rsquo;s own usage is a disclosed internal engine: the team states it gates Copilot code review changes on ReviewBench before A/B tests, which is a concrete claim that the offline signal tracks production.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Cautions&#xA;    &lt;div id=&#34;cautions&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#cautions&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;The publication gate cuts against the leaderboard&amp;rsquo;s meaning: submissions stay private until a maintainer approves them, and scores publish only if they beat the agent&amp;rsquo;s current leaderboard score (or on first entry), so the public board can only ratchet upward and losing runs stay invisible.&lt;/strong&gt;&lt;/li&gt;&#xA;&lt;li&gt;GitHub&amp;rsquo;s team generated the initial commercial entries itself by running the publicly available products, the vendors neither conducted nor verified those runs, and the test dates differ (Copilot on 2026-10-01, Greptile and Cubic back in June, per The New Stack).&lt;/li&gt;&#xA;&lt;li&gt;219 PRs is a small corpus, the augmented-recall denominator moves with each agent&amp;rsquo;s own discoveries and is explicitly not cross-comparable, and the judge is a frontier model whose biases the rubric mitigates but does not remove.&lt;/li&gt;&#xA;&lt;li&gt;Adoption is thin so far: 26 repository stars as of 2026-10-07 and no meaningful independent replication of the leaderboard numbers.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Pricing&#xA;    &lt;div id=&#34;pricing&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#pricing&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;Free to read and free to submit; GitHub provides the judge, and your cost is your own reviewer&amp;rsquo;s model tokens.&#xA;The benchmark datasets are MIT-licensed.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Compared to&#xA;    &lt;div id=&#34;compared-to&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#compared-to&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/evaluation-review/frontierharness-eval/&#34; &gt;FrontierHarness Eval&lt;/a&gt;: the same vendor-runs-the-benchmark structure one object up, scoring coding-agent harnesses instead of reviewers, with its publisher selling agent infrastructure rather than the top entry itself.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/evaluation-review/harnesstax/&#34; &gt;HarnessTax&lt;/a&gt;: the academic counterpart with no product to sell; choose it as the conflict-free cross-examination, ReviewBench for the living leaderboard and self-serve path.&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/evaluation-review/jevals/&#34; &gt;Jevals&lt;/a&gt;: typed decision-model judges for production traces in your own loop; ReviewBench scores reviewer systems offline, before you have picked one.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Bottom line&#xA;    &lt;div id=&#34;bottom-line&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#bottom-line&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;p&gt;&lt;strong&gt;Recommended for teams comparing AI code reviewers who want a reproducible, fully inspectable yardstick and will read the methodology caveats with the leaderboard.&lt;/strong&gt;&#xA;Not for CI gating, and not as an unbiased tiebreaker while the publisher&amp;rsquo;s own reviewer holds the top entry under a publication gate only GitHub controls.&lt;/p&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;Changes&#xA;    &lt;div id=&#34;changes&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#changes&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;2026-10-07 - Created from the daily-refresh entrant resolution, profiling GitHub&amp;rsquo;s open code-review benchmark with the publication-gate and vendor-conflict cautions.&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;See also&#xA;    &lt;div id=&#34;see-also&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#see-also&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/evaluation-review/evaluation-review-feature-matrix/&#34; &gt;Evaluation and Review Feature Matrix&lt;/a&gt; - the category comparison this note joins&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/evaluation-review/frontierharness-eval/&#34; &gt;FrontierHarness Eval&lt;/a&gt; - the vendor-run benchmark precedent for harnesses&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/evaluation-review/harnesstax/&#34; &gt;HarnessTax&lt;/a&gt; - the academic, conflict-free counterweight&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://tomrochette.com/agents/code-review/code-review-feature-matrix/&#34; &gt;Code Review Feature Matrix&lt;/a&gt; - the reviewer products this benchmark scores&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&#xA;&lt;h2 class=&#34;relative group&#34;&gt;References&#xA;    &lt;div id=&#34;references&#34; class=&#34;anchor&#34;&gt;&lt;/div&gt;&#xA;    &#xA;    &lt;span&#xA;        class=&#34;absolute top-0 w-6 transition-opacity opacity-0 -start-6 not-prose group-hover:opacity-100 select-none&#34;&gt;&#xA;        &lt;a class=&#34;text-primary-300 dark:text-neutral-700 !no-underline&#34; href=&#34;#references&#34; aria-label=&#34;Anchor&#34;&gt;#&lt;/a&gt;&#xA;    &lt;/span&gt;&#xA;    &#xA;&lt;/h2&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.blog&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review&lt;/a&gt; - the announcement: corpus, golden set, judge, metrics, submission flow, and the publication gate&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://github.com/review-bench/ReviewBench&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://github.com/review-bench/ReviewBench&lt;/a&gt; - the repository: corpus, mirrors, agent contract, methodology, MIT license&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://raw.githubusercontent.com/review-bench/ReviewBench/main/README.md&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=raw.githubusercontent.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://raw.githubusercontent.com/review-bench/ReviewBench/main/README.md&lt;/a&gt; - corpus composition, test-set tables, mirrors, and the self-serve surface&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://api.github.com/repos/review-bench/ReviewBench&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=api.github.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://api.github.com/repos/review-bench/ReviewBench&lt;/a&gt; - stars, issues, creation and push dates as of 2026-10-07&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://review-bench.ai/&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=review-bench.ai&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://review-bench.ai/&lt;/a&gt; - the leaderboard site (a client-rendered shell to fetchers, so its content is grounded in the announcement and press coverage rather than quoted)&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://thenewstack.io/github-reviewbench-code-review&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=thenewstack.io&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://thenewstack.io/github-reviewbench-code-review&lt;/a&gt; - the critical read: GitHub pre-ran the rival entries, test dates differ, and Copilot&amp;rsquo;s 40.1 grounded F1 leads&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://hn.algolia.com/api/v1/items/49967574&#34;  target=&#34;_blank&#34; rel=&#34;noreferrer&#34;&gt;&lt;img class=&#34;external-link-favicon&#34; src=&#34;https://www.google.com/s2/favicons?domain=hn.algolia.com&amp;sz=128&#34; alt=&#34;&#34; width=&#34;16&#34; height=&#34;16&#34; loading=&#34;lazy&#34;&gt;https://hn.algolia.com/api/v1/items/49967574&lt;/a&gt; - the 4-point, zero-comment launch-day thread, the thin-footprint signal&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;</description>
      
    </item>
    
  </channel>
</rss>
