English
Related papers

Related papers: Benchmarks for Automated Commonsense Reasoning: A …

200 papers

Recently, transformer-based methods such as RoBERTa and GPT-3 have led to significant experimental advances in natural language processing tasks such as question answering and commonsense reasoning. The latter is typically evaluated through…

Computation and Language · Computer Science 2020-11-19 Mayank Kejriwal , Ke Shen

Given the ubiquitous nature of numbers in text, reasoning with numbers to perform simple calculations is an important skill of AI systems. While many datasets and models have been developed to this end, state-of-the-art AI systems are…

Computation and Language · Computer Science 2022-04-13 Swaroop Mishra , Arindam Mitra , Neeraj Varshney , Bhavdeep Sachdeva , Peter Clark , Chitta Baral , Ashwin Kalyan

The primary way to establish and compare competencies in foundation and generative AI models has shifted from peer-reviewed literature to press releases and company blog posts, where model builders highlight results on selected benchmarks.…

Artificial Intelligence · Computer Science 2026-05-15 Stefan Baack , Christo Buschek , Maty Bohacek

Calls for engagement with the public in Artificial Intelligence (AI) research, development, and governance are increasing, leading to the use of surveys to capture people's values, perceptions, and experiences related to AI. In this paper,…

Computers and Society · Computer Science 2024-08-06 Mohammmad Tahaei , Daricia Wilkinson , Alisa Frik , Michael Muller , Ruba Abu-Salma , Lauren Wilcox

Reasoning is not just about solving problems -- it is also about evaluating which problems are worth solving at all. Evaluations of artificial intelligence (AI) systems primarily focused on problem solving, historically by studying how…

Earlier-stage evaluations of a new AI architecture/system need affordable benchmarks. Only using a few AI component benchmarks like MLPerfalone in the other stages may lead to misleading conclusions. Moreover, the learning dynamics are not…

The rapid expansion of large language model (LLM) safety evaluation has produced a substantial benchmark ecosystem, but not a correspondingly coherent measurement ecosystem. We present AISafetyBenchExplorer, a structured catalogue of 195 AI…

Artificial Intelligence · Computer Science 2026-04-24 Abiodun A. Solanke

The pursuit of artificial general intelligence necessitates robust methods for evaluating the cognitive capabilities of models beyond narrow task performance. Here, we introduce a psychometric framework to assess the cognitive profiles of…

Artificial Intelligence · Computer Science 2026-05-11 Isaac Galatzer-Levy , Daniel McDuff , Xin Liu , Jed McGiffin

This position paper argues that the under-representation of social science tasks in contemporary LLM benchmarks limits advances in both LLM evaluation and social scientific inquiry. Benchmarks -- standardized tools for assessing…

Large language models (LLMs) have demonstrated substantial commonsense understanding through numerous benchmark evaluations. However, their understanding of cultural commonsense remains largely unexamined. In this paper, we conduct a…

Computation and Language · Computer Science 2024-05-09 Siqi Shen , Lajanugen Logeswaran , Moontae Lee , Honglak Lee , Soujanya Poria , Rada Mihalcea

Understanding commonsense causality is a unique mark of intelligence for humans. It helps people understand the principles of the real world better and benefits the decision-making process related to causation. For instance, commonsense…

Computation and Language · Computer Science 2024-08-30 Shaobo Cui , Zhijing Jin , Bernhard Schölkopf , Boi Faltings

Today's Internet Services are undergoing fundamental changes and shifting to an intelligent computing era where AI is widely employed to augment services. In this context, many innovative AI algorithms, systems, and architectures are…

With the emergence of advanced reasoning models like OpenAI o3 and DeepSeek-R1, large language models (LLMs) have demonstrated remarkable reasoning capabilities. However, their ability to perform rigorous logical reasoning remains an open…

Artificial Intelligence · Computer Science 2025-02-14 Hanmeng Liu , Zhizhang Fu , Mengru Ding , Ruoxi Ning , Chaoli Zhang , Xiaozhang Liu , Yue Zhang

Researchers are increasingly subjecting artificial intelligence systems to psychological testing. But to rigorously compare their cognitive capacities with humans and other animals, we must avoid both over- and under-stating our…

Artificial Intelligence · Computer Science 2025-03-05 Konstantinos Voudouris , Lucy G. Cheke , Eric Schulz

Recent studies have significantly improved the state-of-the-art on common-sense reasoning (CSR) benchmarks like the Winograd Schema Challenge (WSC) and SWAG. The question we ask in this paper is whether improved performance on these…

Machine Learning · Computer Science 2021-09-27 Paul Trichelair , Ali Emami , Adam Trischler , Kaheer Suleman , Jackie Chi Kit Cheung

The increasing attention on deep learning has tremendously spurred the design of intelligence processing hardware. The variety of emerging intelligence processors requires standard benchmarks for fair comparison and system optimization (in…

Recent advancements in reasoning-reinforced Large Language Models (LLMs) have shown remarkable capabilities in complex reasoning tasks. However, the mechanism underlying their utilization of different human reasoning skills remains poorly…

Computation and Language · Computer Science 2025-08-15 Nghia Trung Ngo , Franck Dernoncourt , Thien Huu Nguyen

Large language models (LLMs) have shown remarkable capabilities in commonsense reasoning; however, some variations in questions can trigger incorrect responses. Do these models truly understand commonsense knowledge, or just memorize…

Computation and Language · Computer Science 2025-05-27 Xiaoyuan Li , Moxin Li , Rui Men , Yichang Zhang , Keqin Bao , Wenjie Wang , Fuli Feng , Dayiheng Liu , Junyang Lin

Conversational Artificial Intelligence (AI) systems have recently sky-rocketed in popularity and are now used in many applications, from car assistants to customer support. The development of conversational AI systems is supported by a…

Human-Computer Interaction · Computer Science 2020-12-23 Johan Aronsson , Philip Lu , Daniel Strüber , Thorsten Berger

Although AI has become increasingly smart, its wisdom has not kept pace. In this article, we examine what is known about human wisdom and sketch a vision of its AI counterpart. We analyze human wisdom as a set of strategies for solving…

‹ Prev 1 4 5 6 7 8 10 Next ›