中文
相关论文

相关论文: Rethinking Theory of Mind Benchmarks for LLMs: Tow…

200 篇论文

Theory of mind (ToM) refers to humans' ability to understand and infer the desires, beliefs, and intentions of others. The acquisition of ToM plays a key role in humans' social cognition and interpersonal relations. Though indispensable for…

计算与语言 · 计算机科学 2024-05-21 Jincenzi Wu , Zhuang Chen , Jiawen Deng , Sahand Sabour , Helen Meng , Minlie Huang

Intrigued by the claims of emergent reasoning capabilities in LLMs trained on general web corpora, in this paper, we set out to investigate their planning capabilities. We aim to evaluate (1) how good LLMs are by themselves in generating…

人工智能 · 计算机科学 2023-02-15 Karthik Valmeekam , Sarath Sreedharan , Matthew Marquez , Alberto Olmo , Subbarao Kambhampati

Cognitive abilities, such as Theory of Mind (ToM), play a vital role in facilitating cooperation in human social interactions. However, our study reveals that agents with higher ToM abilities may not necessarily exhibit better cooperative…

多智能体系统 · 计算机科学 2025-05-15 Jiaqi Shao , Tianjun Yuan , Tao Lin , Bing Luo

Can AI be cognitively biased in automated information judgment tasks? Despite recent progresses in measuring and mitigating social and algorithmic biases in AI and large language models (LLMs), it is not clear to what extent LLMs behave…

信息检索 · 计算机科学 2024-11-26 Jiqun Liu , Jiangen He

Theory-of-Mind (ToM) tasks pose a unique challenge for large language models (LLMs), which often lack the capability for dynamic logical reasoning. In this work, we propose DEL-ToM, a framework that improves verifiable ToM reasoning through…

人工智能 · 计算机科学 2025-09-30 Yuheng Wu , Jianwen Xie , Denghui Zhang , Zhaozhuo Xu

Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive and psychological traits, such as intelligence and personality. While these results are often…

机器学习 · 计算机科学 2026-05-12 Tom Sühr , Florian E. Dorner , Olawale Salaudeen , Augustin Kelava , Samira Samadi

We investigate whether Large Language Models (LLMs) exhibit human-like cognitive patterns under four established frameworks from psychology: Thematic Apperception Test (TAT), Framing Bias, Moral Foundations Theory (MFT), and Cognitive…

人工智能 · 计算机科学 2025-12-12 Akash Kundu , Rishika Goswami

To what extent do LLMs use their capabilities towards their given goal? We take this as a measure of their goal-directedness. We evaluate goal-directedness on tasks that require information gathering, cognitive effort, and plan execution,…

Cognitive biases are systematic deviations in thinking that lead to irrational judgments and problematic decision-making, extensively studied across various fields. Recently, large language models (LLMs) have shown advanced understanding…

计算与语言 · 计算机科学 2024-10-10 Nuo Chen , Jiqun Liu , Xiaoyu Dong , Qijiong Liu , Tetsuya Sakai , Xiao-Ming Wu

Large language models (LLMs) are increasingly used in human-AI interaction research and practice, yet existing capability and safety benchmarks reveal little about the value priorities these systems express or how those priorities…

人工智能 · 计算机科学 2026-05-19 Gabriel Rongyang Lau , Wei Yan Low , Seow Min Koh , Fiona Fui-Hoon Nah , Andree Hartanto

Theory of Mind (ToM) is a fundamental cognitive architecture that endows humans with the ability to attribute mental states to others. Humans infer the desires, beliefs, and intentions of others by observing their behavior and, in turn,…

机器人学 · 计算机科学 2023-11-09 Chuang Yu , Baris Serhan , Angelo Cangelosi

This paper explores the potential of a multidisciplinary approach to testing and aligning artificial intelligence (AI), specifically focusing on large language models (LLMs). Due to the rapid development and wide application of LLMs,…

计算机与社会 · 计算机科学 2025-01-07 Ljubisa Bojic , Matteo Cinelli , Dubravko Culibrk , Boris Delibasic

New developments are enabling AI systems to perceive, recognize, and respond with social cues based on inferences made from humans' explicit or implicit behavioral and verbal cues. These AI systems, equipped with an equivalent of human's…

人机交互 · 计算机科学 2024-05-28 Qiaosi Wang , Ashok K. Goel

In recent years, AI has demonstrated remarkable capabilities in simulating human behaviors, particularly those implemented with large language models (LLMs). However, due to the lack of systematic evaluation of LLMs' simulated behaviors,…

计算与语言 · 计算机科学 2024-06-18 Yang Xiao , Yi Cheng , Jinlan Fu , Jiashuo Wang , Wenjie Li , Pengfei Liu

Large Language Models (LLMs) have shown remarkable capabilities for complex tasks, yet adaptation in medical domain, specifically mental health, poses specific challenges. Mental health is a rising concern globally with LLMs having large…

Peer review is fundamental to scientific research, but the growing volume of publications has intensified the challenges of this expertise-intensive process. While LLMs show promise in various scientific tasks, their potential to assist…

计算与语言 · 计算机科学 2025-07-04 Zhijian Xu , Yilun Zhao , Manasi Patwardhan , Lovekesh Vig , Arman Cohan

The humanlike responses of large language models (LLMs) have prompted social scientists to investigate whether LLMs can be used to simulate human participants in experiments, opinion polls and surveys. Of central interest in this line of…

计算与语言 · 计算机科学 2024-05-14 Nikolay B Petrov , Gregory Serapio-García , Jason Rentfrow

In this paper, we propose a novel personalized decision support system that combines Theory of Mind (ToM) modeling and explainable Reinforcement Learning (XRL) to provide effective and interpretable interventions. Our method leverages DRL…

机器学习 · 计算机科学 2023-12-15 Huao Li , Yao Fan , Keyang Zheng , Michael Lewis , Katia Sycara

With the release of ChatGPT and other large language models (LLMs) the discussion about the intelligence, possibilities, and risks, of current and future models have seen large attention. This discussion included much debated scenarios…

人工智能 · 计算机科学 2024-07-31 Nils Körber , Silvan Wehrli , Christopher Irrgang

Large language models (LLMs) are increasingly used to simulate human behavior in social settings such as legal mediation, negotiation, and dispute resolution. However, it remains unclear whether these simulations reproduce the…

人工智能 · 计算机科学 2026-02-10 Deuksin Kwon , Kaleen Shrestha , Bin Han , Spencer Lin , James Hale , Jonathan Gratch , Maja Matarić , Gale M. Lucas