English

GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models

Artificial Intelligence 2025-10-08 v1 Computation and Language

Abstract

We introduce \textbf{GAUSS} (\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics), a benchmark that evaluates LLMs' mathematical abilities across twelve core skill dimensions, grouped into three domains: knowledge and understanding, problem solving and communication, and meta-skills and creativity. By categorizing problems according to cognitive skills and designing tasks that isolate specific abilities, GAUSS constructs comprehensive, fine-grained, and interpretable profiles of models' mathematical abilities. These profiles faithfully represent their underlying mathematical intelligence. To exemplify how to use the \textsc{GAUSS} benchmark, we have derived the skill profile of \textsc{GPT-5-thinking}, revealing its strengths and weaknesses as well as its differences relative to \textsc{o4-mini-high}, thereby underscoring the value of multidimensional, skill-based evaluation.

Keywords

Cite

@article{arxiv.2509.18122,
  title  = {GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models},
  author = {Yue Zhang and Jiaxin Zhang and Qiuyu Ren and Tahsin Saffat and Xiaoxuan Liu and Zitong Yang and Banghua Zhu and Yi Ma},
  journal= {arXiv preprint arXiv:2509.18122},
  year   = {2025}
}

Comments

120 pages (including appendix)

R2 v1 2026-07-01T05:50:24.095Z