English
Related papers

Related papers: FrontierMath: A Benchmark for Evaluating Advanced …

200 papers

Frontier AI safety policies highlight automation of AI research and development (R&D) by AI agents as an important capability to anticipate. However, there exist few evaluations for AI R&D capabilities, and none that are highly realistic…

Frontier scientific reasoning is rapidly emerging as a key foundation for advancing AI agents in automated scientific discovery. Deep research agents offer a promising approach to this challenge. These models develop robust problem-solving…

Artificial Intelligence · Computer Science 2026-05-27 Tianshi Zheng , Rui Wang , Xiyun Li , Kelvin Kiu Wai Tam , Newt Nguyen Kim Hue Nam , Wei Fan , Yangqiu Song , Tianqing Fang

AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can…

Computers and Society · Computer Science 2024-12-13 Peter Barnett , Lisa Thiergart

Intersectionality is a critical framework that, through inquiry and praxis, allows us to examine how social inequalities persist through domains of structure and discipline. Given AI fairness' raison d'etre of "fairness", we argue that…

Computers and Society · Computer Science 2023-07-24 Anaelia Ovalle , Arjun Subramonian , Vagrant Gautam , Gilbert Gee , Kai-Wei Chang

A key frontier for Multimodal Large Language Models (MLLMs) is the ability to perform deep mathematical and spatial reasoning directly from images, moving beyond their established success in semantic description. Mathematical surface plots…

Artificial Intelligence · Computer Science 2025-09-10 Nilay Pande , Sahiti Yerramilli , Jayant Sravan Tamarapalli , Rynaa Grover

AI agents could accelerate scientific discovery by automating hypothesis formation, experiment design, coding, execution, and analysis, yet existing benchmarks probe narrow skills in simplified settings. To address this gap, we introduce…

Machine learning (ML) has shown promise for tackling combinatorial optimization (CO), but much of the reported progress relies on small-scale, synthetic benchmarks that fail to capture real-world structure and scale. A core limitation is…

Machine Learning · Computer Science 2026-03-11 Shengyu Feng , Weiwei Sun , Shanda Li , Ameet Talwalkar , Yiming Yang

Inequality proving, crucial across diverse scientific and mathematical fields, tests advanced reasoning skills such as discovering tight bounds and strategic theorem application. This makes it a distinct, demanding frontier for large…

Artificial Intelligence · Computer Science 2025-12-16 Pan Lu , Jiayi Sheng , Luna Lyu , Jikai Jin , Tony Xia , Alex Gu , James Zou

Many real-world coding challenges are open-ended and admit no known optimal solution. Yet, recent progress in LLM coding has focused on well-defined tasks such as feature implementation, bug fixing, and competitive programming. Open-ended…

Existing mathematical reasoning benchmarks are predominantly English only or translation-based, which can introduce semantic drift and mask languagespecific reasoning errors. To address this, we present AI4Math, a benchmark of 105 original…

Foundation models (FMs) have achieved significant success across various tasks, leading to research on benchmarks for reasoning abilities. However, there is a lack of studies on FMs performance in exceptional scenarios, which we define as…

Artificial Intelligence · Computer Science 2024-12-06 Suho Kang , Jungyang Park , Joonseo Ha , SoMin Kim , JinHyeong Kim , Subeen Park , Kyungwoo Song

Recent failures such as Google Gemini generating people of color in Nazi-era uniforms illustrate how AI outputs can be factually plausible yet socially harmful. AI models are increasingly evaluated for "fairness," yet existing benchmarks…

Computation and Language · Computer Science 2025-10-01 Jen-tse Huang , Yuhang Yan , Linqi Liu , Yixin Wan , Wenxuan Wang , Kai-Wei Chang , Michael R. Lyu

Recent advancements in large language models (LLMs) have revitalized philosophical debates surrounding artificial intelligence. Two of the most fundamental challenges - namely, the Frame Problem and the Symbol Grounding Problem - have…

Artificial Intelligence · Computer Science 2025-06-10 Shoko Oka

In this study, we explored the progression trajectories of artificial intelligence (AI) systems through the lens of complexity theory. We challenged the conventional linear and exponential projections of AI advancement toward Artificial…

Artificial Intelligence · Computer Science 2024-07-08 Teo Susnjak , Timothy R. McIntosh , Andre L. C. Barczak , Napoleon H. Reyes , Tong Liu , Paul Watters , Malka N. Halgamuge

Tables and table-based use cases play a crucial role in many important real-world applications, such as spreadsheets, databases, and computational notebooks, which traditionally require expert-level users like data engineers, data analysts,…

Artificial Intelligence · Computer Science 2026-03-10 Junjie Xing , Yeye He , Mengyu Zhou , Haoyu Dong , Shi Han , Lingjiao Chen , Dongmei Zhang , Surajit Chaudhuri , H. V. Jagadish

Scientific research communities are embracing AI-based solutions to target tractable scientific tasks and improve research workflows. However, the development and evaluation of such solutions are scattered across multiple disciplines. We…

Artificial Intelligence · Computer Science 2022-06-14 Yatao Li , Jianfeng Zhan

While LLMs have shown impressive capabilities in solving math or coding problems, the ability to make scientific discoveries remains a distinct challenge. This paper proposes a "Turing test for an AI scientist" to assess whether an AI agent…

Artificial Intelligence · Computer Science 2024-05-24 Xiaoxin Yin

We argue how AI can assist mathematics in three ways: theorem-proving, conjecture formulation, and language processing. Inspired by initial experiments in geometry and theoretical physics in 2017, we summarize how this emerging field has…

History and Overview · Mathematics 2025-11-24 Yang-Hui He

Current regulations on powerful AI capabilities are narrowly focused on "foundation" or "frontier" models. However, these terms are vague and inconsistently defined, leading to an unstable foundation for governance efforts. Critically,…

Computers and Society · Computer Science 2024-09-27 Ritwik Gupta , Leah Walker , Rodolfo Corona , Stephanie Fu , Suzanne Petryk , Janet Napolitano , Trevor Darrell , Andrew W. Reddie
‹ Prev 1 8 9 10 Next ›