中文
相关论文

相关论文: Aletheia tackles FirstProof autonomously

200 篇论文

As Large Language Models (LLMs) grow in capability, do they develop self-awareness as an emergent behavior? And if so, can we measure it? We introduce the AI Self-Awareness Index (AISAI), a game-theoretic framework for measuring…

人工智能 · 计算机科学 2025-12-04 Kyung-Hoon Kim

The rapid advancement of LLMs has generated growing interest in their potential role in physics education and assessment, yet a focused evaluation of their performance on multi-faceted, free-response physics problems remains underexplored.…

物理教育 · 物理学 2026-03-10 Bilas Paul , Jashandeep Kaur , Shantanu Chakraborty , Shruti Shrestha

Large Language Models (LLMs) are widely used by students, yet their tendency to provide fast and complete answers may discourage reflection and foster overconfidence. We examined how alternative LLM interaction designs support deeper…

人机交互 · 计算机科学 2026-04-13 Elena Eleftheriou , George Pallis , Marios Constantinides

Can artificial intelligence outperform humans at strategic foresight -- the capacity to form accurate judgments about uncertain, high-stakes outcomes before they unfold? We address this question through a fully prospective prediction…

综合经济学 · 经济学 2026-02-03 Felipe A. Csaszar , Aticus Peterson , Daniel Wilde

We introduce Vibe Reasoning, a human-AI collaborative paradigm for solving complex mathematical problems. Our key insight is that frontier AI models already possess the knowledge required to solve challenging problems -- they simply do not…

人工智能 · 计算机科学 2025-12-23 Jiaao Wu , Xian Zhang , Fan Yang , Yinpeng Dong

The considerable mathematical knowledge encoded by the Flyspeck project is combined with external automated theorem provers (ATPs) and machine-learning premise selection methods trained on the proofs, producing an AI system capable of…

人工智能 · 计算机科学 2021-12-03 Cezary Kaliszyk , Josef Urban

Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks have been proposed to evaluate large language model (LLM) performance on deep research…

人工智能 · 计算机科学 2026-05-29 A. J. Lew , Y. Cao , M. J. Buehler

This study introduces a dedicated model aimed at solving the BRAINTEASER task 9 , a novel challenge designed to assess models lateral thinking capabilities through sentence and word puzzles. Our model demonstrates remarkable efficacy,…

计算与语言 · 计算机科学 2024-03-05 Abdelhak Kelious , Mounir Okirim

Equational Theories Project is a collaborative effort, which explores the validity of certain first-order logic implications of certain kind. The project has been completed but triggered further research. This report investigates how much…

计算机科学中的逻辑 · 计算机科学 2025-08-25 Mikoláš Janota

In this short note, we report and analyze a striking event: OpenAI's large language model o3 has outwitted all students in a university exam on thermodynamics. The thermodynamics exam is a difficult hurdle for most students, where they must…

计算工程、金融与科学 · 计算机科学 2025-06-12 Rebecca Loubet , Pascal Zittlau , Marco Hoffmann , Luisa Vollmer , Sophie Fellenz , Heike Leitte , Fabian Jirasek , Johannes Lenhard , Hans Hasse

Existing AI disclosure mandates in scholarship require that AI assistance be reported but leave transparency philosophically unspecified: they fix the duty without explaining what the duty serves. We argue that ethical inquiry is…

计算机与社会 · 计算机科学 2026-05-19 Michele Loi

Current AI approaches to refugee integration optimize narrow objectives such as employment and fail to capture the cultural, emotional, and ethical dimensions critical for long-term success. We introduce EMPATHIA (Enriched Multimodal…

人工智能 · 计算机科学 2025-08-12 Mohamed Rayan Barhdadi , Mehmet Tuncel , Erchin Serpedin , Hasan Kurban

We present E3, an automated review assistant that augments reviewers and engineering teams by identifying decision-relevant technical concerns in research papers. For each concern, E3 reports its nature, its location, its bearing on the…

计算与语言 · 计算机科学 2026-05-27 Yashwardhan Chaudhuri , Sanyam Jain , Paridhi Mundra

In the summer of 2020 OpenAI released its GPT-3 autoregressive language model to much fanfare. While the model has shown promise on tasks in several areas, it has not always been clear when the results were cherry-picked or when they were…

计算与语言 · 计算机科学 2021-06-29 Curt Kohler , Ron Daniel

Large language models have demonstrated remarkable capabilities across diverse reasoning tasks, yet their performance on algorithmic reasoning remains limited. To handle this limitation, we propose PRIME (Policy-Reinforced Iterative…

计算与语言 · 计算机科学 2026-02-13 Jiawei Xu , Zhenyu Yu , Ziqian Bi , Minh Duc Pham , Xiaoyi Qu , Danyang Zhang

The developments in Generative AI technologies have paved the way for numerous innovations in different fields. Recently, Generative AI has been proposed as a competitor to AES systems in evaluating student essays automatically. Considering…

计算与语言 · 计算机科学 2025-10-20 Enis Oğuz

We introduce FrontierMath, a benchmark of hundreds of original, exceptionally challenging mathematics problems crafted and vetted by expert mathematicians. The questions cover most major branches of modern mathematics -- from…

As AI-driven document understanding and processing tools become increasingly prevalent in real-world applications, the need for rigorous evaluation standards has grown increasingly urgent. Existing benchmarks and evaluations often focus on…

We present Ax-Prover, a multi-agent system for automated theorem proving in Lean that can solve problems across diverse scientific domains and operate either autonomously or collaboratively with human experts. To achieve this, Ax-Prover…

The Alexa Prize program has empowered numerous university students to explore, experiment, and showcase their talents in building conversational agents through challenges like the SocialBot Grand Challenge and the TaskBot Challenge. As…