中文

面向自主数学研究

机器学习 2026-03-09 v3 人工智能 计算与语言 计算机与社会

摘要

近期基础模型的进步已催生出能够在国际数学奥林匹克金 medal 标准水平上达到的推理系统。然而,从竞赛水平的问题解决迈向专业研究,需跨越庞大的文献体系并构建长期证明路径。本文中,我们引入 Aletheia,一款数学研究智能体,能够在自然语言中迭代生成、验证和修订解决方案。具体而言,Aletheia 由高级版 Gemini Deep Think 驱动,能够解决具有挑战性的推理问题,配合超越奥林匹克水平问题的新型推理时序扩展规律,以及大量工具使用以应对数学研究的复杂性。我们展示了 Aletheia 从奥林匹克问题到 PhD 水平练习的能力,最重要的是,通过以下几个独特里程碑实现了 AI 辅助数学研究:(a) 一篇由 AI 完全无人中介生成的研究论文 (Feng26),在算术几何中计算某些结构常数即特征权值的部分;(b) 一篇论文 (LeeSeo26),展示了人类-AI 协作如何证明称为独立集的相互作用粒子系统的界限;以及 (c) 对 Bloom 的 Erdos 猜想数据库中 700 个公开问题的大规模半自主评估,包括四个开放性问题的自主求解。为帮助公众更好地理解 AI 与数学发展相关的进展,我们建议量化 AI 辅助结果的标准自主性与新颖性水平,并提出一种新的人类-AI 交互卡牌概念以提升透明度。我们以反思人类-AI 在数学中的协作,并在 https://github.com/google-deepmind/superhuman/tree/main/aletheia 分享所有提示及模型输出。

关键词

引用

@article{arxiv.2602.10177,
  title  = {Towards Autonomous Mathematics Research},
  author = {Tony Feng and Trieu H. Trinh and Garrett Bingham and Dawsen Hwang and Yuri Chervonyi and Junehyuk Jung and Joonkyung Lee and Carlo Pagano and Sang-hyun Kim and Federico Pasqualotto and Sergei Gukov and Jonathan N. Lee and Junsu Kim and Kaiying Hou and Golnaz Ghiasi and Yi Tay and YaGuang Li and Chenkai Kuang and Yuan Liu and Hanzhao Lin and Evan Zheran Liu and Nigamaa Nayakanti and Xiaomeng Yang and Heng-Tze Cheng and Demis Hassabis and Koray Kavukcuoglu and Quoc V. Le and Thang Luong},
  journal= {arXiv preprint arXiv:2602.10177},
  year   = {2026}
}

备注

42 pages, updated with summary of FirstProof results. Accompanied blog post https://deepmind.google/blog/accelerating-mathematical-and-scientific-discovery-with-gemini-deep-think/