English
Related papers

Related papers: FANToM: A Benchmark for Stress-testing Machine The…

200 papers

The use of LLMs in natural language reasoning has shown mixed results, sometimes rivaling or even surpassing human performance in simpler classification tasks while struggling with social-cognitive reasoning, a domain where humans naturally…

Computation and Language · Computer Science 2024-10-23 Fiona Anting Tan , Gerard Christopher Yeo , Kokil Jaidka , Fanyou Wu , Weijie Xu , Vinija Jain , Aman Chadha , Yang Liu , See-Kiong Ng

Theory of Mind (ToM)-the ability to reason about the mental states of oneself and others-is a cornerstone of human social intelligence. As Large Language Models (LLMs) become ubiquitous in real-world applications, validating their capacity…

Computation and Language · Computer Science 2026-03-13 Ruirui Chen , Weifeng Jiang , Chengwei Qin , Cheston Tan

Theory-of-Mind (ToM) is a fundamental psychological capability that allows humans to understand and interpret the mental states of others. Humans infer others' thoughts by integrating causal cues and indirect clues from broad contextual…

Computation and Language · Computer Science 2025-04-10 Chulun Zhou , Qiujing Wang , Mo Yu , Xiaoqian Yue , Rui Lu , Jiangnan Li , Yifan Zhou , Shunchi Zhang , Jie Zhou , Wai Lam

Existing dynamic Theory of Mind (ToM) benchmarks mostly place language models in a passive role: the model reads a sequence of connected scenarios and reports what people believe, feel, intend, and do as these states change. In real social…

Artificial Intelligence · Computer Science 2026-01-28 Zhichao Liang , Satoshi Nakamura

Large Language Models (LLMs) perform well on many language tasks, but their Theory of Mind (ToM) reasoning is still uneven in complex social settings. Existing benchmarks, including ExploreToM, do not always test the recursive beliefs and…

Recent evidence suggests Large Language Models (LLMs) display Theory of Mind (ToM) abilities. Most ToM experiments place participants in a spectatorial role, wherein they predict and interpret other agents' behavior. However, human ToM also…

Computation and Language · Computer Science 2025-07-23 Jared Moore , Ned Cooper , Rasmus Overmark , Beba Cibralic , Nick Haber , Cameron R. Jones

Most existing Theory of Mind (ToM) benchmarks for foundation models rely on variations of the Sally-Anne test, offering only a very limited perspective on ToM and neglecting the complexity of human social interactions. To address this gap,…

Computation and Language · Computer Science 2025-09-17 Matteo Bortoletto , Constantin Ruhdorfer , Andreas Bulling

We introduce CHARTOM, a visual theory-of-mind benchmark designed to evaluate multimodal large language models' capability to understand and reason about misleading data visualizations though charts. CHARTOM consists of carefully designed…

Artificial Intelligence · Computer Science 2025-07-01 Shubham Bharti , Shiyun Cheng , Jihyun Rho , Jianrui Zhang , Mu Cai , Yong Jae Lee , Martina Rau , Xiaojin Zhu

As Large Language Models (LLMs) gain agentic abilities, they will have to navigate complex multi-agent scenarios, interacting with human users and other agents in cooperative and competitive settings. This will require new reasoning skills,…

Artificial Intelligence · Computer Science 2025-06-26 Andrei Lupu , Timon Willi , Jakob Foerster

Theory of Mind (ToM), the ability to attribute mental states to others, is a hallmark of social intelligence. While large language models (LLMs) demonstrate promising performance on standard ToM benchmarks, we observe that they often fail…

Computation and Language · Computer Science 2026-04-14 Mengfan Li , Xuanhua Shi , Yang Deng

Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and humans. However, the existing benchmarks often measure ToM capability improvement through…

Artificial Intelligence · Computer Science 2026-05-18 Nanxu Gong , Zixin Chen , Haotian Li , Zishu Zhao , Jianxun Lian , Huamin Qu , Yanjie Fu , Xing Xie

Theory of Mind (ToM) assesses whether models can infer hidden mental states such as beliefs, desires, and intentions, which is essential for natural social interaction. Although recent progress in Large Reasoning Models (LRMs) has boosted…

Artificial Intelligence · Computer Science 2026-03-05 Nanxu Gong , Haotian Li , Sixun Dong , Jianxun Lian , Yanjie Fu , Xing Xie

Theory of Mind (ToM) refers to the cognitive ability to infer and attribute mental states to oneself and others. As large language models (LLMs) are increasingly evaluated for social and cognitive capabilities, it remains unclear to what…

Computation and Language · Computer Science 2024-11-26 Jayanta Sadhu , Ayan Antik Khan , Noshin Nawal , Sanju Basak , Abhik Bhattacharjee , Rifat Shahriyar

Theory of Mind (ToM), the capacity to comprehend the mental states of distinct individuals, is essential for numerous practical applications. With the development of large language models (LLMs), there is a heated debate about whether they…

Computation and Language · Computer Science 2024-10-29 Xiaomeng Ma , Lingyu Gao , Qihui Xu

Theory of Mind (ToM) is a critical component of intelligence but its assessment remains the subject of heated debates. Prior research applied human ToM assessments to natural language processing models using either human-created…

Computation and Language · Computer Science 2023-11-08 Damien Sileo , Antoine Lernould

We introduce StorySim, a programmable framework for synthetically generating stories to evaluate the theory of mind (ToM) and world modeling (WM) capabilities of large language models (LLMs). Unlike prior benchmarks that may suffer from…

Computation and Language · Computer Science 2026-04-28 Nathaniel Getachew , Abulhair Saparov

To what degree should we ascribe cognitive capacities to Large Language Models (LLMs), such as the ability to reason about intentions and beliefs known as Theory of Mind (ToM)? Here we add to this emerging debate by (i) testing 11 base- and…

Computation and Language · Computer Science 2023-11-01 Max J. van Duijn , Bram M. A. van Dijk , Tom Kouwenhoven , Werner de Valk , Marco R. Spruit , Peter van der Putten

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D dataset to benchmark the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Yuxuan Li , Vijay Veerabadran , Michael L. Iuzzolino , Brett D. Roads , Asli Celikyilmaz , Karl Ridgeway

Recent studies have increasingly demonstrated that large language models (LLMs) possess significant theory of mind (ToM) capabilities, showing the potential for simulating the tracking of mental states in generative agents. In this study,…

Computation and Language · Computer Science 2025-01-28 Bo Yang , Jiaxian Guo , Yusuke Iwasawa , Yutaka Matsuo

Theory of Mind (ToM)-the cognitive ability to reason about mental states of ourselves and others, is the foundation of social interaction. Although ToM comes naturally to humans, it poses a significant challenge to even the most advanced…

Computation and Language · Computer Science 2024-07-02 Guiyang Hou , Wenqi Zhang , Yongliang Shen , Linjuan Wu , Weiming Lu