English

FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification

Artificial Intelligence 2026-05-28 v3 Machine Learning

Abstract

LLM-based reviewing systems typically take only the manuscript as input, leaving literature and code-based claims hard to verify. We present FactReview, a system that extracts review-relevant claims, grounds them in related work, and, when code is available, executes released artifacts under a fixed repair budget to audit empirical claims. Across 35 ML papers and 463 benchmark major claims, FactReview covers 84% of claims. Under an evidence-aware rubric, its reviews score 4.86/5 in overall quality, 0.7 above DeepReview-v2 and 1.5 above matched OpenReview comments. Removing execution evidence changes 17% of claim statuses, more than any other single evidence source. In a reviewer-assistance study, FactReview reduces mean review time by 58% while raising benchmark claim coverage from 87% to 99%. We argue that LLM reviewers should audit empirical claims, not make accept-reject decisions. The code is public at: https://github.com/DEFENSE-SEU/FactReview.

Keywords

Cite

@article{arxiv.2604.04074,
  title  = {FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification},
  author = {Ling Yue and Chaoqian Ouyang and Hang Xu and Ruijun Huang and Yuchen Liu and Libin Zheng and Wei Liu and Shaowu Pan and Shimin Di and Min-Ling Zhang},
  journal= {arXiv preprint arXiv:2604.04074},
  year   = {2026}
}
R2 v1 2026-07-01T11:54:25.049Z