GAUSS:用于评估大语言模型结构化数学技能的基准
人工智能
2025-10-08 v1 计算与语言
摘要
我们提出\textbf{GAUSS}(\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics),用于评估大语言模型(LLM)在十二个核心技能维度上的数学能力,这些维度分为三个领域:知识与理解、问题解决与沟通,以及元技能与创造力。通过对问题进行认知技能分类并设计隔离特定能力的任务,GAUSS 构建了全面、细致且可解释的数学能力轮廓。这些轮廓忠实地代表了模型的底层数学智能。为展示如何使用 \textsc{GAUSS} 基准,我们对 \textsc{GPT-5-thinking} 进行了技能轮廓分析,揭示了其相对于 \textsc{o4-mini-high} 的优势与不足,进而凸显了多维度、基于技能的评估价值。
引用
@article{arxiv.2509.18122,
title = {GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models},
author = {Yue Zhang and Jiaxin Zhang and Qiuyu Ren and Tahsin Saffat and Xiaoxuan Liu and Zitong Yang and Banghua Zhu and Yi Ma},
journal= {arXiv preprint arXiv:2509.18122},
year = {2025}
}
备注
120 pages (including appendix)