English

Do Repetitions Matter? Strengthening Reliability in LLM Evaluations

Artificial Intelligence 2025-09-30 v1 Computation and Language

Abstract

LLM leaderboards often rely on single stochastic runs, but how many repetitions are required for reliable conclusions remains unclear. We re-evaluate eight state-of-the-art models on the AI4Math Benchmark with three independent runs per setting. Using mixed-effects logistic regression, domain-level marginal means, rank-instability analysis, and run-to-run reliability, we assessed the value of additional repetitions. Our findings shows that Single-run leaderboards are brittle: 10/12 slices (83\%) invert at least one pairwise rank relative to the three-run majority, despite a zero sign-flip rate for pairwise significance and moderate overall interclass correlation. Averaging runs yields modest SE shrinkage (\sim5\% from one to three) but large ranking gains; two runs remove \sim83\% of single-run inversions. We provide cost-aware guidance for practitioners: treat evaluation as an experiment, report uncertainty, and use 2\geq 2 repetitions under stochastic decoding. These practices improve robustness while remaining feasible for small teams and help align model comparisons with real-world reliability.

Keywords

Cite

@article{arxiv.2509.24086,
  title  = {Do Repetitions Matter? Strengthening Reliability in LLM Evaluations},
  author = {Miguel Angel Alvarado Gonzalez and Michelle Bruno Hernandez and Miguel Angel Peñaloza Perez and Bruno Lopez Orozco and Jesus Tadeo Cruz Soto and Sandra Malagon},
  journal= {arXiv preprint arXiv:2509.24086},
  year   = {2025}
}
R2 v1 2026-07-01T06:03:04.786Z