Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark
Abstract
While large language models (LLMs) with reasoning capabilities are progressing rapidly on high-school math competitions and coding, can they reason effectively through complex, open-ended challenges found in frontier physics research? And crucially, what kinds of reasoning tasks do physicists want LLMs to assist with? To address these questions, we present the CritPt (Complex Research using Integrated Thinking - Physics Test, pronounced "critical point"), the first benchmark designed to test LLMs on unpublished, research-level reasoning tasks that broadly covers modern physics research areas, including condensed matter, quantum physics, atomic, molecular & optical physics, astrophysics, high energy physics, mathematical physics, statistical physics, nuclear physics, nonlinear dynamics, fluid dynamics and biophysics. CritPt consists of 71 composite research challenges designed to simulate full-scale research projects at the entry level, which are also decomposed to 190 simpler checkpoint tasks for more fine-grained insights. All problems are newly created by 50+ active physics researchers based on their own research. Every problem is hand-curated to admit a guess-resistant and machine-verifiable answer and is evaluated by an automated grading pipeline heavily customized for advanced physics-specific output formats. We find that while current state-of-the-art LLMs show early promise on isolated checkpoints, they remain far from being able to reliably solve full research-scale challenges: the best average accuracy among base models is only 5.7%, achieved by GPT-5 (high), moderately rising to around 10% when equipped with coding tools. Through the realistic yet standardized evaluation offered by CritPt, we highlight a large disconnect between current model capabilities and realistic physics research demands, offering a foundation to guide the development of scientifically grounded AI tools.
Keywords
Cite
@article{arxiv.2509.26574,
title = {Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark},
author = {Minhui Zhu and Minyang Tian and Xiaocheng Yang and Tianci Zhou and Lifan Yuan and Penghao Zhu and Eli Chertkov and Shengyan Liu and Yufeng Du and Ziming Ji and Indranil Das and Qingzhi Chen and Junyi Cao and Yufeng Du and Jiabin Yu and Peixue Wu and Jinchen He and Yifan Su and Yikun Jiang and Yujie Zhang and Chang Liu and Ze-Min Huang and Weizhen Jia and Yunkai Wang and Farshid Jafarpour and Yong Zhao and Xinan Chen and Jessie Shelton and Aaron W. Young and John Bartolotta and Wenchao Xu and Yue Sun and Anjun Chu and Victor Colussi and Chris Akers and Nathan Brooks and Wenbo Fu and Jinchao Zhao and Marvin Qi and Anqi Mu and Yubo Yang and Allen Zang and Yang Lyu and Peizhi Mai and Christopher Wilson and Xuefei Guo and Juntai Zhou and Daniel Inafuku and Chi Xue and Luyu Gao and Ze Yang and Yaïr Hein and Yonatan Kahn and Kevin Zhou and Di Luo and John Drew Wilson and Jarrod T. Reilly and Dmytro Bandak and Ofir Press and Liang Yang and Xueying Wang and Hao Tong and Nicolas Chia and Eliu Huerta and Hao Peng},
journal= {arXiv preprint arXiv:2509.26574},
year = {2026}
}
Comments
40 pages, 6 figures, 6 tables