English
Related papers

Related papers: Evaluating Theory of Mind in Question Answering

200 papers

Reasoning models have achieved remarkable performance on tasks like math and logical reasoning thanks to their ability to search during reasoning. However, they still suffer from overthinking, often performing unnecessary reasoning steps…

Artificial Intelligence · Computer Science 2025-04-09 Anqi Zhang , Yulin Chen , Jane Pan , Chen Zhao , Aurojit Panda , Jinyang Li , He He

Can emergent language models faithfully model the intelligence of decision-making agents? Though modern language models exhibit already some reasoning ability, and theoretically can potentially express any probable distribution over tokens,…

Machine Learning · Computer Science 2024-06-27 Wenhao Lu , Xufeng Zhao , Josua Spisak , Jae Hee Lee , Stefan Wermter

Scaling the test-time compute of large language models has demonstrated impressive performance on reasoning benchmarks. However, existing evaluations of test-time scaling make the strong assumption that a reasoning system should always give…

Computation and Language · Computer Science 2025-07-21 William Jurayj , Jeffrey Cheng , Benjamin Van Durme

Test-time scaling increases inference-time computation by allowing models to generate long reasoning chains, and has improved performance across many domains. However, in this work, we show that this approach is not yet effective for…

Artificial Intelligence · Computer Science 2026-02-03 James Xu Zhao , Bryan Hooi , See-Kiong Ng

Large language models (LLMs) are capable of generating plausible explanations of how they arrived at an answer to a question. However, these explanations can misrepresent the model's "reasoning" process, i.e., they can be unfaithful. This,…

Computation and Language · Computer Science 2025-05-21 Katie Matton , Robert Osazuwa Ness , John Guttag , Emre Kıcıman

A growing body of work makes use of probing to investigate the working of neural models, often considered black boxes. Recently, an ongoing debate emerged surrounding the limitations of the probing paradigm. In this work, we point out the…

Computation and Language · Computer Science 2021-02-22 Yanai Elazar , Shauli Ravfogel , Alon Jacovi , Yoav Goldberg

Many studies have evaluated the cognitive alignment of Pre-trained Language Models (PLMs), i.e., their correspondence to adult performance across a range of cognitive domains. Recently, the focus has expanded to the developmental alignment…

Computation and Language · Computer Science 2025-01-23 Raj Sanjay Shah , Sashank Varma

When people think of everyday things like an egg, they typically have a mental image associated with it. This allows them to correctly judge, for example, that "the yolk surrounds the shell" is a false statement. Do language models…

Computation and Language · Computer Science 2023-06-09 Yuling Gu , Bhavana Dalvi Mishra , Peter Clark

As artificial agents become increasingly capable, what internal structure is *necessary* for an agent to act competently under uncertainty? Classical results show that optimal control can be *implemented* using belief states or world…

Machine Learning · Computer Science 2026-04-03 Aran Nayebi

Human interactions are deeply rooted in the interplay of thoughts, beliefs, and desires made possible by Theory of Mind (ToM): our cognitive ability to understand the mental states of ourselves and others. Although ToM may come naturally to…

Artificial Intelligence · Computer Science 2023-11-20 Alex Wilf , Sihyun Shawn Lee , Paul Pu Liang , Louis-Philippe Morency

Present language understanding methods have demonstrated extraordinary ability of recognizing patterns in texts via machine learning. However, existing methods indiscriminately use the recognized patterns in the testing phase that is…

Computation and Language · Computer Science 2021-06-08 Fuli Feng , Jizhi Zhang , Xiangnan He , Hanwang Zhang , Tat-Seng Chua

Training language models with rationales augmentation has been shown to be beneficial in many existing works. In this paper, we identify that such a prevailing view does not hold consistently. We conduct comprehensive investigations to…

Computation and Language · Computer Science 2025-06-02 Chiwei Zhu , Benfeng Xu , An Yang , Junyang Lin , Quan Wang , Chang Zhou , Zhendong Mao

With their recent development, large language models (LLMs) have been found to exhibit a certain level of Theory of Mind (ToM), a complex cognitive capacity that is related to our conscious mind and that allows us to infer another's beliefs…

Computation and Language · Computer Science 2023-09-06 Mohsen Jamali , Ziv M. Williams , Jing Cai

The impressive performance of language models is undeniable. However, the presence of biases based on gender, race, socio-economic status, physical appearance, and sexual orientation makes the deployment of language models challenging. This…

Computation and Language · Computer Science 2025-08-13 Swati Rajwal , Shivank Garg , Reem Abdel-Salam , Abdelrahman Zayed

Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision-making, and existing benchmarks focus solely on human…

Artificial Intelligence · Computer Science 2026-02-25 Ruoxuan Zhang , Qiyun Zheng , Zhiyu Zhou , Ziqi Liao , Siyu Wu , Jian-Yu Jiang-Lin , Bin Wen , Hongxia Xie , Jianlong Fu , Wen-Huang Cheng

Mathematical reasoning---a core ability within human intelligence---presents some unique challenges as a domain: we do not come to understand and solve mathematical problems primarily on the back of experience and evidence, but on the basis…

Machine Learning · Computer Science 2019-04-03 David Saxton , Edward Grefenstette , Felix Hill , Pushmeet Kohli

When we test a theory using data, it is common to focus on correctness: do the predictions of the theory match what we see in the data? But we also care about completeness: how much of the predictable variation in the data is captured by…

Machine Learning · Computer Science 2017-06-22 Jon Kleinberg , Annie Liang , Sendhil Mullainathan

I propose that pattern recognition, memorization and processing are key concepts that can be a principle set for the theoretical modeling of the mind function. Most of the questions about the mind functioning can be answered by a…

Artificial Intelligence · Computer Science 2009-07-28 Gilberto de Paiva

World Models help Artificial Intelligence (AI) predict outcomes, reason about its environment, and guide decision-making. While widely used in reinforcement learning, they lack the structured, adaptive representations that even young…

Artificial Intelligence · Computer Science 2025-03-20 Javier Del Ser , Jesus L. Lobo , Heimo Müller , Andreas Holzinger

Motivated reasoning - the idea that individuals processing information may be motivated to either arrive at accurate beliefs or arrive at desired conclusions - has been well-explored as a human phenomenon. However, it remains unclear…

Human-Computer Interaction · Computer Science 2026-05-11 Neeley Pate , Adiba Mahbub Proma , Hangfeng He , James N. Druckman , Daniel C. Molden , Gourab Ghoshal , Ehsan Hoque