English

Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions

Computation and Language 2025-08-14 v4 Artificial Intelligence

Abstract

Recent studies have raised significant concerns regarding the reliability of current mathematics benchmarks, highlighting issues such as simplistic design and potential data contamination. Consequently, developing a reliable benchmark that effectively evaluates large language models' (LLMs) genuine capabilities in mathematical reasoning remains a critical challenge. To address these concerns, we propose RV-Bench, a novel evaluation methodology for Benchmarking LLMs with Random Variables in mathematical reasoning. Specifically, we build question-generating functions to produce random variable questions (RVQs), whose background content mirrors original benchmark problems, but with randomized variable combinations, rendering them "unseen" to LLMs. Models must completely understand the inherent question pattern to correctly answer RVQs with diverse variable combinations. Thus, an LLM's genuine reasoning capability is reflected through its accuracy and robustness on RV-Bench. We conducted extensive experiments on over 30 representative LLMs across more than 1,000 RVQs. Our findings propose that LLMs exhibit a proficiency imbalance between encountered and ``unseen'' data distributions. Furthermore, RV-Bench reveals that proficiency generalization across similar mathematical reasoning tasks is limited, but we verified it can still be effectively elicited through test-time scaling.

Keywords

Cite

@article{arxiv.2501.11790,
  title  = {Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions},
  author = {Zijin Hong and Hao Wu and Su Dong and Junnan Dong and Yilin Xiao and Yujing Zhang and Zhu Wang and Feiran Huang and Linyi Li and Hongxia Yang and Xiao Huang},
  journal= {arXiv preprint arXiv:2501.11790},
  year   = {2025}
}
R2 v1 2026-06-28T21:11:52.668Z