English

U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs

Computation and Language 2026-02-03 v4 Artificial Intelligence

Abstract

The current evaluation of mathematical skills in LLMs is limited, as existing benchmarks are either relatively small, primarily focus on elementary and high-school problems, or lack diversity in topics. Additionally, the inclusion of visual elements in tasks remains largely under-explored. To address these gaps, we introduce U-MATH, a novel benchmark of 1,100 unpublished open-ended university-level problems sourced from teaching materials. It is balanced across six core subjects, with 20% of multimodal problems. Given the open-ended nature of U-MATH problems, we employ an LLM to judge the correctness of generated solutions. To this end, we release μ\mu-MATH, a dataset to evaluate the LLMs' capabilities in judging solutions. Benchmarking leading LLMs reveals marked limitations in multi-modal reasoning, with maximum accuracy reaching 93.1\% on textual tasks but only 58.5\% on visual ones. Furthermore, solution judgment proves challenging, requiring the most advanced models to achieve meaningfully high performance, even still peaking at an imperfect F1-score of 90.1\%.

Keywords

Cite

@article{arxiv.2412.03205,
  title  = {U-MATH: A University-Level Benchmark for Evaluating Mathematical Skills in LLMs},
  author = {Konstantin Chernyshev and Vitaliy Polshkov and Ekaterina Artemova and Alex Myasnikov and Vlad Stepanov and Alexei Miasnikov and Sergei Tilga},
  journal= {arXiv preprint arXiv:2412.03205},
  year   = {2026}
}
R2 v1 2026-06-28T20:22:45.486Z