English

Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks

Software Engineering 2026-02-10 v4 Artificial Intelligence Computation and Language

Abstract

Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the community interprets model capabilities. In the past few years, awareness of benchmark quality has grown. Yet, after a decade-scale (2014-2025) survey over 572 code benchmarks, we observed a lag between growing awareness and actual practice. For example, in 2025 alone, the number of benchmarks that ignore code coverage when providing test cases nearly matches the total count accumulated across the previous ten years. In response, we take a clear position: Code benchmarks must prioritize rigor in benchmark construction, reliability in evaluation, and reproducibility in release. To operationalize this position, we introduce a code benchmark guideline HOW2BENCH with 55 checklists. Finally, our further human study also exposed that the current issues not only stem from the significant effort required, but also from a lack of awareness regarding their importance.

Keywords

Cite

@article{arxiv.2501.10711,
  title  = {Rigor, Reliability, and Reproducibility Matter: A Decade-Scale Survey of 572 Code Benchmarks},
  author = {Jialun Cao and Yuk-Kit Chan and Zixuan Ling and Wenxuan Wang and Shuqing Li and Mingwei Liu and Ruixi Qiao and Yuting Han and Chaozheng Wang and Boxi Yu and Pinjia He and Shuai Wang and Zibin Zheng and Michael R. Lyu and Shing-Chi Cheung},
  journal= {arXiv preprint arXiv:2501.10711},
  year   = {2026}
}

Comments

65 pages