English
Related papers

Related papers: FrontierScience: Evaluating AI's Ability to Perfor…

200 papers

AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions of inquiry; indeed, there are now many such agents, ranging…

Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However, existing benchmarks rarely test a fundamental bottleneck: whether Large Language…

Machine Learning · Computer Science 2026-05-29 Sy-Tuyen Ho , Minghui Liu , Huy Nghiem , Furong Huang

Reasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmarks often fail to rigorously evaluate the nuanced reasoning…

Autonomous scientific research is significantly advanced thanks to the development of AI agents. One key step in this process is finding the right scientific literature, whether to explore existing knowledge for a research problem, or to…

Frontier AI models -- highly capable foundation models at the cutting edge of AI development -- may pose severe risks to public safety, human rights, economic stability, and societal value in the coming years. These risks could arise from…

Computers and Society · Computer Science 2025-03-11 Deepika Raman , Nada Madkour , Evan R. Murphy , Krystal Jackson , Jessica Newman

AI-assisted research is crossing a threshold: fully automated systems can now generate research papers for as little as $15, while long-horizon agents can execute experiments, draft manuscripts, and simulate critique with minimal human…

Since language models (LMs) now outperform average humans on many challenging tasks, it has become increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this issue by examining LMs' capabilities…

Frontier AI agents show increasing promise as scientific research assistants, and may eventually be useful for extended, open-ended research workflows. However, in order to use agents for novel research, we must first assess the underlying…

AI research agents have shown strong potential in automating literature search and manuscript refinement, yet most assume a clear and actionable initial input, operating only after a research question has been made explicit. In contrast,…

Artificial Intelligence · Computer Science 2026-05-08 Jie Yu , Song Qiu

We present a fully reproducible demonstration of an AI-assisted scientific workflow designed for a broad physics, mathematics, and computer-science readership. The initial project artifact stack was generated from one single user prompt and…

Other Condensed Matter · Physics 2026-03-17 Kin Hung Fung

Large language models (LLMs) are playing an increasingly important role in scientific research, yet there remains a lack of comprehensive benchmarks to evaluate the breadth and depth of scientific knowledge embedded in these models. To…

Computation and Language · Computer Science 2025-10-08 Kehua Feng , Xinyi Shen , Weijie Wang , Xiang Zhuang , Yuqi Tang , Qiang Zhang , Keyan Ding

We introduce SciEvalKit, a unified benchmarking toolkit designed to evaluate AI models for science across a broad range of scientific disciplines and task capabilities. Unlike general-purpose evaluation platforms, SciEvalKit focuses on the…

Recent advances in foundational models have yielded reasoning systems capable of achieving a gold-medal standard at the International Mathematical Olympiad. The transition from competition-level problem-solving to professional research,…

Large language models (LLMs) are increasingly being used for complex research tasks such as literature review, idea generation, and scientific paper analysis, yet their ability to truly understand and process the intricate relationships…

Computation and Language · Computer Science 2025-06-11 Shashidhar Reddy Javaji , Yupeng Cao , Haohang Li , Yangyang Yu , Nikhil Muralidhar , Zining Zhu

Not yet. We present SPACE, a benchmark that systematically evaluates spatial cognition in frontier models. Our benchmark builds on decades of research in cognitive science. It evaluates large-scale mapping abilities that are brought to bear…

Artificial Intelligence · Computer Science 2025-04-21 Santhosh Kumar Ramakrishnan , Erik Wijmans , Philipp Kraehenbuehl , Vladlen Koltun

As large language models (LLMs) transition from general knowledge retrieval to complex scientific discovery, their evaluation standards must also incorporate the rigorous norms of scientific inquiry. Existing benchmarks exhibit a critical…

Artificial Intelligence · Computer Science 2026-01-13 Encheng Su , Jianyu Wu , Chen Tang , Lintao Wang , Pengze Li , Aoran Wang , Jinouwen Zhang , Yizhou Wang , Yuan Meng , Xinzhu Ma , Shixiang Tang , Houqiang Li

Advanced AI models hold the promise of tremendous benefits for humanity, but society needs to proactively manage the accompanying risks. In this paper, we focus on what we term "frontier AI" models: highly capable foundation models that…

Recent advances in machine learning and AI, including Generative AI and LLMs, are disrupting technological innovation, product development, and society as a whole. AI's contribution to technology can come from multiple approaches that…

In the upcoming decade, deep learning may revolutionize the natural sciences, enhancing our capacity to model and predict natural occurrences. This could herald a new era of scientific exploration, bringing significant advancements across…

Artificial Intelligence · Computer Science 2023-10-13 Shuaiwen Leon Song , Bonnie Kruft , Minjia Zhang , Conglong Li , Shiyang Chen , Chengming Zhang , Masahiro Tanaka , Xiaoxia Wu , Jeff Rasley , Ammar Ahmad Awan , Connor Holmes , Martin Cai , Adam Ghanem , Zhongzhu Zhou , Yuxiong He , Pete Luferenko , Divya Kumar , Jonathan Weyn , Ruixiong Zhang , Sylwester Klocek , Volodymyr Vragov , Mohammed AlQuraishi , Gustaf Ahdritz , Christina Floristean , Cristina Negri , Rao Kotamarthi , Venkatram Vishwanath , Arvind Ramanathan , Sam Foreman , Kyle Hippe , Troy Arcomano , Romit Maulik , Maxim Zvyagin , Alexander Brace , Bin Zhang , Cindy Orozco Bohorquez , Austin Clyde , Bharat Kale , Danilo Perez-Rivera , Heng Ma , Carla M. Mann , Michael Irvin , J. Gregory Pauloski , Logan Ward , Valerie Hayot , Murali Emani , Zhen Xie , Diangen Lin , Maulik Shukla , Ian Foster , James J. Davis , Michael E. Papka , Thomas Brettin , Prasanna Balaprakash , Gina Tourassi , John Gounley , Heidi Hanson , Thomas E Potok , Massimiliano Lupo Pasini , Kate Evans , Dan Lu , Dalton Lunga , Junqi Yin , Sajal Dash , Feiyi Wang , Mallikarjun Shankar , Isaac Lyngaas , Xiao Wang , Guojing Cong , Pei Zhang , Ming Fan , Siyan Liu , Adolfy Hoisie , Shinjae Yoo , Yihui Ren , William Tang , Kyle Felker , Alexey Svyatkovskiy , Hang Liu , Ashwin Aji , Angela Dalton , Michael Schulte , Karl Schulz , Yuntian Deng , Weili Nie , Josh Romero , Christian Dallago , Arash Vahdat , Chaowei Xiao , Thomas Gibbs , Anima Anandkumar , Rick Stevens

Given rapid progress toward advanced AI and risks from frontier AI systems (advanced AI systems pushing the boundaries of the AI capabilities frontier), the creation and implementation of AI governance and regulatory schemes deserves…