中文
相关论文

相关论文: Re-evaluating Theory of Mind evaluation in large l…

200 篇论文

While recent studies explore Large Language Models' (LLMs) performance on Theory of Mind (ToM) reasoning tasks, research on ToM abilities that require more nuanced social context is limited, such as white lies. We introduce TactfulToM, a…

计算与语言 · 计算机科学 2025-09-26 Yiwei Liu , Emma Jane Pretty , Jiahao Huang , Saku Sugawara

Theory of Mind (ToM) is central to social cognition and human-AI interaction, and Large Language Models (LLMs) have been used to help understand and represent ToM. However, most evaluations treat ToM as a static judgment at a single moment,…

人工智能 · 计算机科学 2026-03-17 Thuy Ngoc Nguyen , Duy Nhat Phan , Cleotilde Gonzalez

This paper examines the extent to which large language models (LLMs) have developed higher-order theory of mind (ToM); the human ability to reason about multiple mental and emotional states in a recursive manner (e.g. I think that you…

Improving the Theory of Mind (ToM) capability of Large Language Models (LLMs) is crucial for effective social interactions between these AI models and humans. However, the existing benchmarks often measure ToM capability improvement through…

人工智能 · 计算机科学 2026-05-18 Nanxu Gong , Zixin Chen , Haotian Li , Zishu Zhao , Jianxun Lian , Huamin Qu , Yanjie Fu , Xing Xie

Theory of Mind (ToM) is the ability to attribute mental states to others, the basis of human cognition. At present, there has been growing interest in the AI with cognitive abilities, for example in healthcare and the motoring industry.…

人工智能 · 计算机科学 2023-03-22 Yuanyuan Mao , Shuang Liu , Pengshuai Zhao , Qin Ni , Xin Lin , Liang He

Reasoning is a fundamental aspect of human intelligence that plays a crucial role in activities such as problem solving, decision making, and critical thinking. In recent years, large language models (LLMs) have made significant progress in…

计算与语言 · 计算机科学 2023-05-29 Jie Huang , Kevin Chen-Chuan Chang

We propose a hybrid approach to machine Theory of Mind (ToM) that uses large language models (LLMs) as a mechanism for generating hypotheses and likelihood functions with a Bayesian inverse planning model that computes posterior…

人工智能 · 计算机科学 2025-07-08 Rebekah A. Gelpí , Eric Xue , William A. Cunningham

With the success of ChatGPT and other similarly sized SotA LLMs, claims of emergent human like social reasoning capabilities, especially Theory of Mind (ToM), in these models have appeared in the scientific literature. On the one hand those…

计算与语言 · 计算机科学 2024-10-10 Christian Nickel , Laura Schrewe , Lucie Flek

Datasets used for emotion recognition tasks typically contain overt cues that can be used in predicting the emotions expressed in a text. However, one challenge is that texts sometimes contain covert contextual cues that are rich in…

计算与语言 · 计算机科学 2025-06-03 Gerard Christopher Yeo , Kokil Jaidka

Theory of Mind (ToM), the ability to attribute mental states to others, is a hallmark of social intelligence. While large language models (LLMs) demonstrate promising performance on standard ToM benchmarks, we observe that they often fail…

计算与语言 · 计算机科学 2026-04-14 Mengfan Li , Xuanhua Shi , Yang Deng

Whether Large Language Models (LLMs) truly possess human-like Theory of Mind (ToM) capabilities has garnered increasing attention. However, existing benchmarks remain largely restricted to narrow paradigms like false belief tasks, failing…

人工智能 · 计算机科学 2026-01-23 Haibo Tong , Zeyang Yue , Feifei Zhao , Erliang Lin , Lu Jia , Ruolin Chen , Yinqian Sun , Qian Zhang , Yi Zeng

Intuitive psychology is a pillar of common-sense reasoning. The replication of this reasoning in machine intelligence is an important stepping-stone on the way to human-like artificial intelligence. Several recent tasks and benchmarks for…

人工智能 · 计算机科学 2023-03-15 Tomer Ullman

As large language models evolve, there is growing anticipation that they will emulate human-like Theory of Mind (ToM) to assist with routine tasks. However, existing methods for evaluating machine ToM focus primarily on unimodal models and…

人工智能 · 计算机科学 2025-06-18 Xinyang Li , Siqi Liu , Bochao Zou , Jiansheng Chen , Huimin Ma

Recent advancements in large language models (LLMs) have demonstrated emergent capabilities in complex reasoning, largely spurred by rule-based Reinforcement Learning (RL) techniques applied during the post-training. This has raised the…

机器学习 · 计算机科学 2025-07-22 Sneheel Sarangi , Hanan Salam

This study explored how large language models (LLMs) perform in two areas related to art: writing critiques of artworks and reasoning about mental states (Theory of Mind, or ToM) in art-related situations. For the critique generation part,…

计算与语言 · 计算机科学 2025-09-16 Takaya Arita , Wenxian Zheng , Reiji Suzuki , Fuminori Akiba

Large Language models are revolutionizing the conversational recommender systems through their impressive capabilities in instruction comprehension, reasoning, and human interaction. A core factor underlying effective recommendation…

人工智能 · 计算机科学 2025-12-01 Mengfan Li , Xuanhua Shi , Yang Deng

Theory of Mind (ToM) capabilities in LLMs have recently become a central object of investigation. Cognitive science distinguishes between two steps required for ToM tasks: 1) determine whether to invoke ToM, which includes the appropriate…

人工智能 · 计算机科学 2025-06-03 Eitan Wagner , Nitay Alon , Joseph M. Barnby , Omri Abend

Our paper argues that the majority of theory of mind benchmarks are broken because of their inability to directly test how large language models (LLMs) adapt to new partners. This problem stems from the fact that theory of mind benchmarks…

人工智能 · 计算机科学 2025-06-13 Matthew Riemer , Zahra Ashktorab , Djallel Bouneffouf , Payel Das , Miao Liu , Justin D. Weisz , Murray Campbell

Natural language interaction with agentic Artificial Intelligence (AI), driven by Large Language Models (LLMs), is expected to remain a dominant paradigm in the near future. While humans instinctively align their communication with mental…

计算与语言 · 计算机科学 2025-05-21 Mehdi Jafari , Devin Yuncheng Hua , Hao Xue , Flora Salim

Theory of Mind (ToM) is a critical component of intelligence but its assessment remains the subject of heated debates. Prior research applied human ToM assessments to natural language processing models using either human-created…

计算与语言 · 计算机科学 2023-11-08 Damien Sileo , Antoine Lernould