中文
相关论文

相关论文: Re-evaluating Theory of Mind evaluation in large l…

200 篇论文

As Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial. While existing benchmarks…

计算与语言 · 计算机科学 2025-06-10 Yang Xiao , Jiashuo Wang , Qiancheng Xu , Changhe Song , Chunpu Xu , Yi Cheng , Wenjie Li , Pengfei Liu

Neural Theory-of-Mind (N-ToM), machine's ability to understand and keep track of the mental states of others, is pivotal in developing socially intelligent agents. However, prevalent N-ToM benchmarks have several shortcomings, including the…

人工智能 · 计算机科学 2024-06-04 Hainiu Xu , Runcong Zhao , Lixing Zhu , Jinhua Du , Yulan He

A growing body of work attempts to evaluate the theory of mind (ToM) abilities of humans and large language models (LLMs) using static, non-interactive question-and-answer benchmarks. However, theoretical work in the field suggests that…

计算与语言 · 计算机科学 2026-02-20 Jared Moore , Rasmus Overmark , Ned Cooper , Beba Cibralic , Nick Haber , Cameron R. Jones

Recent studies have increasingly demonstrated that large language models (LLMs) possess significant theory of mind (ToM) capabilities, showing the potential for simulating the tracking of mental states in generative agents. In this study,…

计算与语言 · 计算机科学 2025-01-28 Bo Yang , Jiaxian Guo , Yusuke Iwasawa , Yutaka Matsuo

The cognitive mechanism by which Large Language Models (LLMs) solve mathematical problems remains a widely debated and unresolved issue. Currently, there is little interpretable experimental evidence that connects LLMs' problem-solving with…

人工智能 · 计算机科学 2025-09-23 Wei Xie , Shuoyoucheng Ma , Zhenhua Wang , Enze Wang , Kai Chen , Xiaobing Sun , Baosheng Wang

Theory of Mind (ToM)-an understanding of the mental states of others-is a key aspect of human social intelligence, yet, chatbots and LLM-based social agents do not typically integrate it. In this work, we demonstrate that LLMs that…

计算与语言 · 计算机科学 2026-04-14 EunJeong Hwang , Yuwei Yin , Giuseppe Carenini , Peter West , Vered Shwartz

Large language models (LLMs) have recently shown impressive performance on tasks involving reasoning, leading to a lively debate on whether these models possess reasoning capabilities similar to humans. However, despite these successes, the…

计算与语言 · 计算机科学 2024-08-07 Philipp Mondorf , Barbara Plank

Theory of Mind (ToM)-the cognitive ability to reason about mental states of ourselves and others, is the foundation of social interaction. Although ToM comes naturally to humans, it poses a significant challenge to even the most advanced…

计算与语言 · 计算机科学 2024-07-02 Guiyang Hou , Wenqi Zhang , Yongliang Shen , Linjuan Wu , Weiming Lu

Theory-of-Mind (ToM), the ability to infer others' perceptions and mental states, is fundamental to human interaction but remains challenging for Large Language Models (LLMs). While existing ToM reasoning methods show promise with reasoning…

计算与语言 · 计算机科学 2025-06-03 Hainiu Xu , Siya Qi , Jiazheng Li , Yuxiang Zhou , Jinhua Du , Caroline Catmur , Yulan He

Large Language Models (LLMs) have shown potential in simulating human behaviors and performing theory-of-mind (ToM) reasoning, a crucial skill for complex social interactions. In this study, we investigate the role of ToM reasoning in…

计算与语言 · 计算机科学 2025-06-02 Neemesh Yadav , Palakorn Achananuparp , Jing Jiang , Ee-Peng Lim

Safety fine-tuning in Large Language Models (LLMs) seeks to suppress potentially harmful forms of mind-attribution such as models asserting their own consciousness or claiming to experience emotions. We investigate whether suppressing…

计算与语言 · 计算机科学 2026-04-01 Junsol Kim , Winnie Street , Roberta Rocca , Daine M. Korngiebel , Adam Waytz , James Evans , Geoff Keeling

Current Large Language Models (LLMs) are unparalleled in their ability to generate grammatically correct, fluent text. LLMs are appearing rapidly, and debates on LLM capacities have taken off, but reflection is lagging behind. Thus, in this…

计算与语言 · 计算机科学 2023-11-01 Bram M. A. van Dijk , Tom Kouwenhoven , Marco R. Spruit , Max J. van Duijn

Large Language Models (LLMs) excel in generating personalized content and facilitating interactive dialogues, showcasing their remarkable aptitude for a myriad of applications. However, their capabilities in reasoning and providing…

计算与语言 · 计算机科学 2024-02-16 Min Zhang , Sato Takumi , Jack Zhang , Jun Wang

This paper investigates whether LMs recruit shared computational mechanisms for general Theory of Mind (ToM) and language-specific pragmatic reasoning in order to contribute to the general question of whether LMs may be said to have…

计算与语言 · 计算机科学 2026-04-28 Polina Tsvilodub , Jan-Felix Klumpp , Amir Mohammadpour , Jennifer Hu , Michael Franke

The performance of Large language models (LLMs) across a broad range of domains has been impressive but have been critiqued as not being able to reason about their process and conclusions derived. This is to explain the conclusions draw,…

计算与语言 · 计算机科学 2024-10-30 Rob Sullivan , Nelly Elsayed

Do large language models (LLMs) display rational reasoning? LLMs have been shown to contain human biases due to the data they have been trained on; whether this is reflected in rational reasoning remains less clear. In this paper, we answer…

计算与语言 · 计算机科学 2024-02-16 Olivia Macmillan-Scott , Mirco Musolesi

In the present study, we investigate and compare reasoning in large language models (LLM) and humans using a selection of cognitive psychology tools traditionally dedicated to the study of (bounded) rationality. To do so, we presented to…

计算与语言 · 计算机科学 2023-09-25 Nicolas Yax , Hernan Anlló , Stefano Palminteri

Integrated Information Theory (IIT) provides a quantitative framework for explaining consciousness phenomenon, positing that conscious systems comprise elements integrated through causal properties. We apply IIT 3.0 and 4.0 -- the latest…

计算与语言 · 计算机科学 2025-07-01 Jingkai Li

How might messages about large language models (LLMs) found in public discourse influence the way people think about and interact with these models? To explore this question, we randomly assigned participants (N = 470) to watch short…

Theory of Mind (ToM), the ability to understand people's minds based on their behavior, is key to developing socially intelligent agents. Current approaches to ToM reasoning either rely on prompting Large Language Models (LLMs), which are…

人工智能 · 计算机科学 2026-01-15 Zhining Zhang , Chuanyang Jin , Mung Yao Jia , Shunchi Zhang , Tianmin Shu