English
Related papers

Related papers: AGIEval: A Human-Centric Benchmark for Evaluating …

200 papers

This article presents the results and their discussion for the third wave (with n=23 participants) within a multinational longitudinal study that investigates the evolving paradigm of human-AI collaboration in problem-solving contexts.…

Computers and Society · Computer Science 2025-12-16 Matthias Huemmer , Theophile Shyiramunda , Franziska Durner , Michelle J. Cummings-Koether

Human understanding and generation are critical for modeling digital humans and humanoid embodiments. Recently, Human-centric Foundation Models (HcFMs) inspired by the success of generalist models, such as large language and vision models,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Shixiang Tang , Yizhou Wang , Lu Chen , Yuan Wang , Sida Peng , Dan Xu , Wanli Ouyang

Large language models (LLMs) have made significant progress in natural language processing tasks and demonstrate considerable potential in the legal domain. However, legal applications demand high standards of accuracy, reliability, and…

Computation and Language · Computer Science 2024-11-27 Haitao Li , You Chen , Qingyao Ai , Yueyue Wu , Ruizhe Zhang , Yiqun Liu

The evaluation of mathematical reasoning capabilities is essential for advancing Artificial General Intelligence (AGI). While Large Language Models (LLMs) have shown impressive performance in solving mathematical problems, existing…

Computation and Language · Computer Science 2025-01-15 Bo Yang , Qingping Yang , Yingwei Ma , Runtao Liu

As general-purpose artificial intelligence systems become increasingly integrated into society and are used for information seeking, content generation, problem solving, textual analysis, coding, and running processes, it is crucial to…

Computers and Society · Computer Science 2025-08-28 Ljubisa Bojic , Dylan Seychell , Milan Cabarkapa

Reasoning, a crucial ability for complex problem-solving, plays a pivotal role in various real-world settings such as negotiation, medical diagnosis, and criminal investigation. It serves as a fundamental methodology in the field of…

The Abstraction and Reasoning Corpus (ARC) poses a stringent test of general AI capabilities, requiring solvers to infer abstract patterns from only a handful of examples. Despite substantial progress in deep learning, state-of-the-art…

Artificial Intelligence · Computer Science 2025-05-28 Woochang Sim , Hyunseok Ryu , Kyungmin Choi , Sungwon Han , Sundong Kim

Recent breakthroughs in diffusion models, multimodal pretraining, and efficient finetuning have led to an explosion of text-to-image generative models. Given human evaluation is expensive and difficult to scale, automated methods are…

Computer Vision and Pattern Recognition · Computer Science 2023-10-19 Dhruba Ghosh , Hanna Hajishirzi , Ludwig Schmidt

We introduce CFE-Bench (Classroom Final Exam), a multimodal benchmark for evaluating the reasoning capabilities of large language models across more than 20 STEM domains. CFE-Bench is curated from repeatedly used, authentic university…

Artificial Intelligence · Computer Science 2026-03-04 Chongyang Gao , Diji Yang , Shuyan Zhou , Xichen Yan , Luchuan Song , Shuo Li , Kezhen Chen

We argue that a key reasoning skill that any advanced AI, say GPT-4, should master in order to qualify as 'thinking machine', or AGI, is hypothetic-deductive reasoning. Problem-solving or question-answering can quite generally be construed…

Artificial Intelligence · Computer Science 2023-08-08 Louis Vervoort , Vitaliy Mizyakov , Anastasia Ugleva

Citation quality is crucial in information-seeking systems, directly influencing trust and the effectiveness of information access. Current evaluation frameworks, both human and automatic, mainly rely on Natural Language Inference (NLI) to…

Computation and Language · Computer Science 2025-06-03 Yumo Xu , Peng Qi , Jifan Chen , Kunlun Liu , Rujun Han , Lan Liu , Bonan Min , Vittorio Castelli , Arshit Gupta , Zhiguo Wang

We present and evaluate a suite of proof-of-concept (PoC), structured workflow prompts designed to elicit human-like hierarchical reasoning while guiding Large Language Models (LLMs) in the high-level semantic and linguistic analysis of…

Computation and Language · Computer Science 2025-06-18 Evgeny Markhasin

Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm drains computational and annotation resources, inflates…

Computation and Language · Computer Science 2026-02-03 Shuai Zhang , Jiayu Hu , Zijie Chen , Zeyuan Ding , Yi Zhang , Yingji Zhang , Ziyi Zhou , Junwei Liao , Shengjie Zhou , Yong Dai , Zhenzhong Lan , Xiaozhu Ju

Facing the current debate on whether Large Language Models (LLMs) attain near-human intelligence levels (Mitchell & Krakauer, 2023; Bubeck et al., 2023; Kosinski, 2023; Shiffrin & Mitchell, 2023; Ullman, 2023), the current study introduces…

Artificial Intelligence · Computer Science 2024-05-21 Junqi Wang , Chunhui Zhang , Jiapeng Li , Yuxi Ma , Lixing Niu , Jiaheng Han , Yujia Peng , Yixin Zhu , Lifeng Fan

As foundation models (FMs) approach human-level fluency, distinguishing synthetic from organic content has become a key challenge for Trustworthy Web Intelligence. This paper presents JudgeGPT and RogueGPT, a dual-axis framework that…

Computers and Society · Computer Science 2026-02-13 Alexander Loth , Martin Kappes , Marc-Oliver Pahl

In the summer of 2020 OpenAI released its GPT-3 autoregressive language model to much fanfare. While the model has shown promise on tasks in several areas, it has not always been clear when the results were cherry-picked or when they were…

Computation and Language · Computer Science 2021-06-29 Curt Kohler , Ron Daniel

We present FormalProofBench, a private benchmark designed to evaluate whether AI models can produce formally verified mathematical proofs at the graduate level. Each task pairs a natural-language problem with a Lean~4 formal statement, and…

Artificial Intelligence · Computer Science 2026-03-31 Nikil Ravi , Kexing Ying , Vasilii Nesterov , Rayan Krishnan , Elif Uskuplu , Bingyu Xia , Janitha Aswedige , Langston Nashold

Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks have been proposed to evaluate large language model (LLM) performance on deep research…

Artificial Intelligence · Computer Science 2026-05-29 A. J. Lew , Y. Cao , M. J. Buehler

Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Professional, an open benchmark for evaluating large language…

Deep research -- producing comprehensive, citation-grounded reports by searching and synthesizing information from hundreds of live web sources -- marks an important frontier for agentic systems. To rigorously evaluate this ability, four…

Artificial Intelligence · Computer Science 2026-04-21 Jiayu Wang , Yifei Ming , Riya Dulepet , Qinglin Chen , Austin Xu , Zixuan Ke , Frederic Sala , Aws Albarghouthi , Caiming Xiong , Shafiq Joty