中文
相关论文

相关论文: RecToM: A Benchmark for Evaluating Machine Theory …

200 篇论文

Datasets used for emotion recognition tasks typically contain overt cues that can be used in predicting the emotions expressed in a text. However, one challenge is that texts sometimes contain covert contextual cues that are rich in…

计算与语言 · 计算机科学 2025-06-03 Gerard Christopher Yeo , Kokil Jaidka

Theory of Mind (ToM) refers to the cognitive ability to infer and attribute mental states to oneself and others. As large language models (LLMs) are increasingly evaluated for social and cognitive capabilities, it remains unclear to what…

计算与语言 · 计算机科学 2024-11-26 Jayanta Sadhu , Ayan Antik Khan , Noshin Nawal , Sanju Basak , Abhik Bhattacharjee , Rifat Shahriyar

In recent years, evaluating the Theory of Mind (ToM) capabilities of large language models (LLMs) has received significant attention within the research community. As the field rapidly evolves, navigating the diverse approaches and…

计算与语言 · 计算机科学 2025-02-14 Karahan Sarıtaş , Kıvanç Tezören , Yavuz Durmazkeser

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D dataset to benchmark the…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Yuxuan Li , Vijay Veerabadran , Michael L. Iuzzolino , Brett D. Roads , Asli Celikyilmaz , Karl Ridgeway

Do Large Language Models (LLMs) possess a Theory of Mind (ToM)? Research into this question has focused on evaluating LLMs against benchmarks and found success across a range of social tasks. However, these evaluations do not test for the…

人工智能 · 计算机科学 2026-02-27 John Muchovej , Amanda Royka , Shane Lee , Julian Jara-Ettinger

Theory of Mind (ToM) refers to the ability of individuals to attribute mental states to others. While Large Language Models (LLMs) have shown some promise with ToM ability, they still struggle with complex ToM reasoning. Our approach…

计算与语言 · 计算机科学 2024-06-27 Weizhi Tang , Vaishak Belle

Theory of Mind (ToM) refers to an agent's ability to model the internal states of others. Contributing to the debate whether large language models (LLMs) exhibit genuine ToM capabilities, our study investigates their ToM robustness using…

计算与语言 · 计算机科学 2026-02-26 Christian Nickel , Laura Schrewe , Florian Mai , Lucie Flek

Natural language interaction with agentic Artificial Intelligence (AI), driven by Large Language Models (LLMs), is expected to remain a dominant paradigm in the near future. While humans instinctively align their communication with mental…

计算与语言 · 计算机科学 2025-05-21 Mehdi Jafari , Devin Yuncheng Hua , Hao Xue , Flora Salim

Understanding and attributing mental states, known as Theory of Mind (ToM), emerges as a fundamental capability for human social reasoning. While Large Language Models (LLMs) appear to possess certain ToM abilities, the mechanisms…

人工智能 · 计算机科学 2024-05-31 Wentao Zhu , Zhining Zhang , Yizhou Wang

We propose a hybrid approach to machine Theory of Mind (ToM) that uses large language models (LLMs) as a mechanism for generating hypotheses and likelihood functions with a Bayesian inverse planning model that computes posterior…

人工智能 · 计算机科学 2025-07-08 Rebekah A. Gelpí , Eric Xue , William A. Cunningham

Theory-of-Mind (ToM) is a fundamental psychological capability that allows humans to understand and interpret the mental states of others. Humans infer others' thoughts by integrating causal cues and indirect clues from broad contextual…

计算与语言 · 计算机科学 2025-04-10 Chulun Zhou , Qiujing Wang , Mo Yu , Xiaoqian Yue , Rui Lu , Jiangnan Li , Yifan Zhou , Shunchi Zhang , Jie Zhou , Wai Lam

Large language models (LLMs) perform substantially below human level on existing theory-of-mind (ToM) benchmarks, even when augmented with chain-of-thought prompting or probabilistic belief updates. We argue that these failures primarily…

计算与语言 · 计算机科学 2026-04-21 Wang Bill Zhu , Qiutong Tony Yi , Robin Jia , Jesse Thomason

Large language models (LLMs) have shown potential in recommendation systems (RecSys) by using them as either knowledge enhancer or zero-shot ranker. A key challenge lies in the large semantic gap between LLMs and RecSys where the former…

信息检索 · 计算机科学 2025-12-22 Guangneng Hu

Theory of Mind (ToM) is central to social cognition and human-AI interaction, and Large Language Models (LLMs) have been used to help understand and represent ToM. However, most evaluations treat ToM as a static judgment at a single moment,…

人工智能 · 计算机科学 2026-03-17 Thuy Ngoc Nguyen , Duy Nhat Phan , Cleotilde Gonzalez

Theory of Mind (ToM)$\unicode{x2014}$the ability to reason about the mental states of other people$\unicode{x2014}$is a key element of our social intelligence. Yet, despite their ever more impressive performance, large-scale neural language…

计算与语言 · 计算机科学 2023-06-02 Melanie Sclar , Sachin Kumar , Peter West , Alane Suhr , Yejin Choi , Yulia Tsvetkov

Recent advancements have showcased the potential of Large Language Models (LLMs) in executing reasoning tasks, particularly facilitated by Chain-of-Thought (CoT) prompting. While tasks like arithmetic reasoning involve clear, definitive…

Theory of Mind (ToM), the ability to attribute mental states to others and predict their behaviour, is fundamental to social intelligence. In this paper, we survey studies evaluating behavioural and representational ToM in Large Language…

计算与语言 · 计算机科学 2025-02-11 Hieu Minh "Jord" Nguyen

Theory of Mind (ToM), the ability to understand the mental states of oneself and others, remains a challenging area for large language models (LLMs), which often fail to predict human mental states accurately. In this paper, we introduce…

We introduce StorySim, a programmable framework for synthetically generating stories to evaluate the theory of mind (ToM) and world modeling (WM) capabilities of large language models (LLMs). Unlike prior benchmarks that may suffer from…

计算与语言 · 计算机科学 2026-04-28 Nathaniel Getachew , Abulhair Saparov

Humans continuously infer the states, goals, and behaviors of others by perceiving their surroundings in dynamic, real-world social interactions. However, most Theory of Mind (ToM) benchmarks only evaluate static, text-based scenarios,…

计算与语言 · 计算机科学 2025-12-16 Xianzhe Fan , Xuhui Zhou , Chuanyang Jin , Kolby Nottingham , Hao Zhu , Maarten Sap