English
Related papers

Related papers: Rethinking Theory of Mind Benchmarks for LLMs: Tow…

200 papers

Job interview simulation with a virtual agents aims at improving people's social skills and supporting professional inclusion. In such simulators, the virtual agent must be capable of representing and reasoning about the user's mental state…

Artificial Intelligence · Computer Science 2014-02-21 Marwen Belkaid , Nicolas Sabouret

Theory of Mind (ToM), the ability to attribute mental states to others, is a hallmark of social intelligence. While large language models (LLMs) demonstrate promising performance on standard ToM benchmarks, we observe that they often fail…

Computation and Language · Computer Science 2026-04-14 Mengfan Li , Xuanhua Shi , Yang Deng

The emergence of instruction-tuned large language models (LLMs) has advanced the field of dialogue systems, enabling both realistic user simulations and robust multi-turn conversational agents. However, existing research often evaluates…

Computation and Language · Computer Science 2025-07-22 Chalamalasetti Kranti , Sherzod Hakimov , David Schlangen

We introduce DialToM, an annotated Theory of Mind (ToM) benchmark built from naturalistic human-human dialogues using a multiple-choice evaluation framework. Concurrent with recent work showing a gap between explicit mental-state inference…

Computation and Language · Computer Science 2026-05-29 Neemesh Yadav , Palakorn Achananuparp , Jing Jiang , Ee-Peng Lim

Relating explicit psychological mechanisms and observable behaviours is a central aim of psychological and behavioural science. We implemented the principles of the Projective Consciousness Model into artificial agents embodied as virtual…

Neurons and Cognition · Quantitative Biology 2025-11-26 David Rudrauf , Grégoire Sergeant-Perthuis , Yvain Tisserand , Teerawat Monnor , Olivier Belli

Theory of Mind (ToM)$\unicode{x2014}$the ability to reason about the mental states of other people$\unicode{x2014}$is a key element of our social intelligence. Yet, despite their ever more impressive performance, large-scale neural language…

Computation and Language · Computer Science 2023-06-02 Melanie Sclar , Sachin Kumar , Peter West , Alane Suhr , Yejin Choi , Yulia Tsvetkov

Existing approaches to Theory of Mind (ToM) in Artificial Intelligence (AI) overemphasize prompted, or cue-based, ToM, which may limit our collective ability to develop Artificial Social Intelligence (ASI). Drawing from research in computer…

Artificial Intelligence · Computer Science 2024-02-22 Nikolos Gurney , David V. Pynadath , Volkan Ustun

Theory of Mind (ToM) is a hallmark of human cognition, allowing individuals to reason about others' beliefs and intentions. Engineers behind recent advances in Artificial Intelligence (AI) have claimed to demonstrate comparable…

Artificial Intelligence · Computer Science 2025-04-01 Nitay Alon , Joseph Barnby , Reuth Mirsky , Stefan Sarkadi

Theory of Mind (ToM) - the ability to attribute beliefs and intents to others - is fundamental for social intelligence, yet Vision-Language Model (VLM) evaluations remain largely Western-centric. In this work, we introduce CulturalToM-VQA,…

Computation and Language · Computer Science 2026-01-08 Zabir Al Nazi , GM Shahariar , Md. Abrar Hossain , Wei Peng

Human social interactions depend on the ability to infer others' unspoken intentions, emotions, and beliefs-a cognitive skill grounded in the psychological concept of Theory of Mind (ToM). While large language models (LLMs) excel in…

Computation and Language · Computer Science 2025-10-15 Xuanming Zhang , Yuxuan Chen , Samuel Yeh , Sharon Li

The validity of online behavioral research relies on study participants being human rather than machine. In the past, it was possible to detect machines by posing simple challenges that were easily solved by humans but not by machines.…

Computation and Language · Computer Science 2026-04-02 Simon Schug , Brenden M. Lake

Large Language Models (LLMs) have recently shown a promise and emergence of Theory of Mind (ToM) ability and even outperform humans in certain ToM tasks. To evaluate and extend the boundaries of the ToM reasoning ability of LLMs, we propose…

Artificial Intelligence · Computer Science 2024-06-10 Weizhi Tang , Vaishak Belle

As large language models (LLMs) are increasingly used to model and augment collective decision-making, it is critical to examine their alignment with human social reasoning. We present an empirical framework for assessing collective…

Artificial Intelligence · Computer Science 2025-10-03 Crystal Qian , Aaron Parisi , Clémentine Bouleau , Vivian Tsai , Maël Lebreton , Lucas Dixon

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to experience emotions. We investigate whether suppressing…

Computation and Language · Computer Science 2026-04-01 Junsol Kim , Winnie Street , Roberta Rocca , Daine M. Korngiebel , Adam Waytz , James Evans , Geoff Keeling

As evaluation designs of large language models may shape our trajectory toward artificial general intelligence, comprehensive and forward-looking assessment is essential. Existing benchmarks primarily assess static knowledge, while…

Computation and Language · Computer Science 2025-08-07 Jiayin Wang , Zhiquang Guo , Weizhi Ma , Min Zhang

Being able to predict the mental states of others is a key factor to effective social interaction. It is also crucial for distributed multi-agent systems, where agents are required to communicate and cooperate. In this paper, we introduce…

Multiagent Systems · Computer Science 2022-04-26 Yuanfei Wang , Fangwei Zhong , Jing Xu , Yizhou Wang

Large language models (LLMs) are increasingly tested for a "Theory of Mind" (ToM) - the ability to attribute mental states to oneself and others. Yet most evaluations stop at explicit belief attribution in classical toy stories or stylized…

Computation and Language · Computer Science 2026-03-03 Yuling Gu , Oyvind Tafjord , Hyunwoo Kim , Jared Moore , Ronan Le Bras , Peter Clark , Yejin Choi

Reasoning is a distinctive human-like characteristic attributed to LLMs in HCI due to their ability to simulate various human-level tasks. However, this work argues that the reasoning behavior of LLMs in HCI is often decontextualized from…

Human-Computer Interaction · Computer Science 2025-10-28 Ramaravind Kommiya Mothilal , Sally Zhang , Syed Ishtiaque Ahmed , Shion Guha

Recent advances in large language models (LLMs) have enabled the emergence of general-purpose agents for automating end-to-end machine learning (ML) workflows, including data analysis, feature engineering, model training, and competition…

Artificial Intelligence · Computer Science 2025-09-12 Hangyi Jia , Yuxi Qian , Hanwen Tong , Xinhui Wu , Lin Chen , Feng Wei

Large language models (LLMs) have advanced conversational AI assistants. However, systematically evaluating how well these assistants apply personalization--adapting to individual user preferences while completing tasks--remains…

Computation and Language · Computer Science 2025-06-12 Zheng Zhao , Clara Vania , Subhradeep Kayal , Naila Khan , Shay B. Cohen , Emine Yilmaz