English
Related papers

Related papers: RiddleSense: Reasoning about Riddle Questions Feat…

200 papers

Current multimodal benchmarks often conflate reasoning with domain-specific knowledge, making it difficult to isolate and evaluate general reasoning abilities in non-expert settings. To address this, we introduce VisualPuzzles, a benchmark…

Computation and Language · Computer Science 2025-05-01 Yueqi Song , Tianyue Ou , Yibo Kong , Zecheng Li , Graham Neubig , Xiang Yue

We propose GuessBench, a novel benchmark that evaluates Vision Language Models (VLMs) on modeling the pervasive, noisy, and pluralistic human creativity. GuessBench sources data from "Guess the Build", an online multiplayer Minecraft…

Computation and Language · Computer Science 2025-06-09 Zifeng Zhu , Shangbin Feng , Herun Wan , Ningnan Wang , Minnan Luo , Yulia Tsvetkov

Representation and learning of commonsense knowledge is one of the foundational problems in the quest to enable deep language understanding. This issue is particularly challenging for understanding casual and correlational relationships…

Computation and Language · Computer Science 2016-04-07 Nasrin Mostafazadeh , Nathanael Chambers , Xiaodong He , Devi Parikh , Dhruv Batra , Lucy Vanderwende , Pushmeet Kohli , James Allen

Visual Dialog requires an agent to engage in a conversation with humans grounded in an image. Many studies on Visual Dialog focus on the understanding of the dialog history or the content of an image, while a considerable amount of…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Shunyu Zhang , Xiaoze Jiang , Zequn Yang , Tao Wan , Zengchang Qin

Commonsense norms are defeasible by context: reading books is usually great, but not when driving a car. While contexts can be explicitly described in language, in embodied scenarios, contexts are often provided visually. This type of…

Machine Learning · Computer Science 2023-11-14 Seungju Han , Junhyeok Kim , Jack Hessel , Liwei Jiang , Jiwan Chung , Yejin Son , Yejin Choi , Youngjae Yu

Large language models (LLMs) have demonstrated potential in reasoning tasks, but their performance on linguistics puzzles remains consistently poor. These puzzles, often derived from Linguistics Olympiad (LO) contests, provide a minimal…

While significant work has been done in the field of NLP on vertical thinking, which involves primarily logical thinking, little work has been done towards lateral thinking, which involves looking at problems from an unconventional…

Computation and Language · Computer Science 2024-05-21 Suyash Vardhan Mathur , Akshett Rai Jindal , Manish Shrivastava

Humans are increasingly interacting with machines through language, sometimes in contexts where the user may not know they are talking to a machine (like over the phone or a text chatbot). We aim to understand how system designers and…

Computation and Language · Computer Science 2021-06-08 David Gros , Yu Li , Zhou Yu

We introduce a new dataset for joint reasoning about natural language and images, with a focus on semantic diversity, compositionality, and visual reasoning challenges. The data contains 107,292 examples of English sentences paired with web…

Computation and Language · Computer Science 2019-07-23 Alane Suhr , Stephanie Zhou , Ally Zhang , Iris Zhang , Huajun Bai , Yoav Artzi

We introduce Drivelology, a unique linguistic phenomenon characterised as "nonsense with depth" - utterances that are syntactically coherent yet pragmatically paradoxical, emotionally loaded, or rhetorically subversive. While such…

Computation and Language · Computer Science 2025-10-17 Yang Wang , Chenghao Xiao , Chia-Yi Hsiao , Zi Yan Chang , Chi-Li Chen , Tyler Loakman , Chenghua Lin

Nowadays, it's a very significant way for researchers and other individuals to achieve their interests because it provides short solutions to satisfy their demands. Because there are so many pieces of information on the internet, news…

Information Retrieval · Computer Science 2022-09-14 Niran A. Abdulhussein , Ahmed J Obaid

We approach the question "What is Consciousness?" in a new way, not as Descartes' "systematic doubt", but as how organisms find their way in their world. Finding one's way involves finding possible uses of features of the world that might…

Physics and Society · Physics 2022-06-30 Stuart A. Kauffman , Andrea Roli

Abductive reasoning aims to find plausible explanations for an event. This style of reasoning is critical for commonsense tasks where there are often multiple plausible explanations. Existing approaches for abductive reasoning in natural…

Computation and Language · Computer Science 2023-05-25 Wenting Zhao , Justin T. Chiu , Claire Cardie , Alexander M. Rush

Recent work by Zellers et al. (2018) introduced a new task of commonsense natural language inference: given an event description such as "A woman sits at a piano," a machine must select the most likely followup: "She sets her fingers on the…

Computation and Language · Computer Science 2019-05-21 Rowan Zellers , Ari Holtzman , Yonatan Bisk , Ali Farhadi , Yejin Choi

Crowdsourcing can solve problems that current fully automated systems cannot. Its effectiveness depends on the reliability, accuracy, and speed of the crowd workers that drive it. These objectives are frequently at odds with one another.…

Human-Computer Interaction · Computer Science 2014-08-29 Walter S. Lasecki , Christopher M. Homan , Jeffrey P. Bigham

We introduce Pencil Puzzle Bench, a framework for evaluating large language model reasoning through pencil puzzles, a family of constraint-satisfaction problems closely related to NP-complete problems, with deterministic, step-level…

Artificial Intelligence · Computer Science 2026-03-03 Justin Waugh

Current commonsense reasoning research focuses on developing models that use commonsense knowledge to answer multiple-choice questions. However, systems designed to answer multiple-choice questions may not be useful in applications that do…

Computation and Language · Computer Science 2021-06-08 Bill Yuchen Lin , Haitian Sun , Bhuwan Dhingra , Manzil Zaheer , Xiang Ren , William W. Cohen

While existing benchmarks probe the reasoning abilities of large language models (LLMs) across diverse domains, they predominantly assess passive reasoning, providing models with all the information needed to reach a solution. By contrast,…

Machine Learning · Computer Science 2025-06-11 Zhanke Zhou , Xiao Feng , Zhaocheng Zhu , Jiangchao Yao , Sanmi Koyejo , Bo Han

This Article introduces the generative reasonable person, a new tool for estimating how ordinary people judge reasonableness. As claims about AI capabilities often outpace evidence, the Article proceeds empirically: adapting randomized…

Computers and Society · Computer Science 2026-02-18 Yonathan A. Arbel

An essential element of human mathematical reasoning is our number sense -- an abstract understanding of numbers and their relationships -- which allows us to solve problems involving vast number spaces using limited computational…

Artificial Intelligence · Computer Science 2025-04-02 Roussel Rahman