English

Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning

Computation and Language 2025-05-22 v1 Machine Learning

Abstract

Large language models (LLMs) demonstrate considerable potential in various natural language tasks but face significant challenges in mathematical reasoning, particularly in executing precise, multi-step logic. However, current evaluation frameworks judge their performance solely based on accuracy, which only accounts for the final answer. This study explores these pitfalls by employing a novel evaluation framework. We propose an evaluation metric called the MAPLE score, which holistically quantifies reasoning misalignment by integrating error rates, redundancy, and validity.

Keywords

Cite

@article{arxiv.2505.15623,
  title  = {Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning},
  author = {Tiasa Singha Roy and Aditeya Baral and Ayush Rajesh Jhaveri and Yusuf Baig},
  journal= {arXiv preprint arXiv:2505.15623},
  year   = {2025}
}
R2 v1 2026-07-01T02:28:53.690Z