中文
相关论文

相关论文: Rethinking Theory of Mind Benchmarks for LLMs: Tow…

200 篇论文

A growing body of work attempts to evaluate the theory of mind (ToM) abilities of humans and large language models (LLMs) using static, non-interactive question-and-answer benchmarks. However, theoretical work in the field suggests that…

计算与语言 · 计算机科学 2026-02-20 Jared Moore , Rasmus Overmark , Ned Cooper , Beba Cibralic , Nick Haber , Cameron R. Jones

The advancement of large language models (LLMs) has outpaced traditional evaluation methodologies. This progress presents novel challenges, such as measuring human-like psychological constructs, moving beyond static and task-specific…

计算与语言 · 计算机科学 2026-03-12 Haoran Ye , Jing Jin , Yuhang Xie , Xin Zhang , Guojie Song

Achieving human-like perception and reasoning in Multimodal Large Language Models (MLLMs) remains a central challenge in artificial intelligence. While recent research has primarily focused on enhancing reasoning capabilities in MLLMs, a…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Hongcheng Gao , Zihao Huang , Lin Xu , Jingyi Tang , Xinhao Li , Yue Liu , Haoyang Li , Taihang Hu , Minhua Lin , Xinlong Yang , Ge Wu , Balong Bi , Hongyu Chen , Wentao Zhang

From pre-trained language model (PLM) to large language model (LLM), the field of natural language processing (NLP) has witnessed steep performance gains and wide practical uses. The evaluation of a research field guides its direction of…

计算与语言 · 计算机科学 2023-08-16 Ziyu Zhuang , Qiguang Chen , Longxuan Ma , Mingda Li , Yi Han , Yushan Qian , Haopeng Bai , Zixian Feng , Weinan Zhang , Ting Liu

The ability of Large Language Models (LLMs) to use external tools unlocks powerful real-world interactions, making rigorous evaluation essential. However, current benchmarks primarily report final accuracy, revealing what models can do but…

计算与语言 · 计算机科学 2026-01-29 Qihao Wang , Yue Hu , Mingzhe Lu , Jiayue Wu , Yanbing Liu , Yuanmin Tang

The concept of persona, originally adopted in dialogue literature, has re-surged as a promising framework for tailoring large language models (LLMs) to specific context (e.g., personalized search, LLM-as-a-judge). However, the growing…

计算与语言 · 计算机科学 2024-10-08 Yu-Min Tseng , Yu-Chao Huang , Teng-Yun Hsiao , Wei-Lin Chen , Chao-Wei Huang , Yu Meng , Yun-Nung Chen

This paper investigates the ability of large language models (LLMs) to solve statistical tasks, as well as their capacity to assess the quality of reasoning. While state-of-the-art LLMs have demonstrated remarkable performance in a range of…

计算与语言 · 计算机科学 2026-01-22 Crish Nagarkar , Leonid Bogachev , Serge Sharoff

Theory of Mind (ToM), the ability to attribute beliefs, intentions, or mental states to others, is a crucial feature of human social interaction. In complex environments, where the human sensory system reaches its limits, behaviour is…

神经与进化计算 · 计算机科学 2024-07-26 Francesca Bianco , Silvia Rigato , Maria Laura Filippetti , Dimitri Ognibene

This research paper delves into the evolving landscape of fine-tuning large language models (LLMs) to align with human users, extending beyond basic alignment to propose "personality alignment" for language models in organizational…

人机交互 · 计算机科学 2023-12-07 Byunggu Yu , Junwhan Kim

Large language models (LLMs) have shown impressive capabilities across various natural language tasks. However, evaluating their alignment with human preferences remains a challenge. To this end, we propose a comprehensive human evaluation…

As large language models (LLMs) become increasingly integrated into society, their alignment with human morals is crucial. To better understand this alignment, we created a large corpus of human- and LLM-generated responses to various moral…

人机交互 · 计算机科学 2024-10-11 Basile Garcia , Crystal Qian , Stefano Palminteri

Currently, in the study of multiagent systems, the intentions of agents are usually ignored. Nonetheless, as pointed out by Theory of Mind (ToM), people regularly reason about other's mental states, including beliefs, goals, and intentions,…

多智能体系统 · 计算机科学 2021-10-04 Luyao Yuan , Zipeng Fu , Linqi Zhou , Kexin Yang , Song-Chun Zhu

We introduce CHARTOM, a visual theory-of-mind benchmark designed to evaluate multimodal large language models' capability to understand and reason about misleading data visualizations though charts. CHARTOM consists of carefully designed…

人工智能 · 计算机科学 2025-07-01 Shubham Bharti , Shiyun Cheng , Jihyun Rho , Jianrui Zhang , Mu Cai , Yong Jae Lee , Martina Rau , Xiaojin Zhu

The paper discusses what is needed to address the limitations of current LLM-centered AI systems. The paper argues that incorporating insights from human cognition and psychology, as embodied by a computational cognitive architecture, can…

人工智能 · 计算机科学 2024-01-22 Ron Sun

Multi-round incomplete information tasks are crucial for evaluating the lateral thinking capabilities of large language models (LLMs). Currently, research primarily relies on multiple benchmarks and automated evaluation metrics to assess…

计算与语言 · 计算机科学 2025-06-02 Wenhan Dong , Tianyi Hu , Jingyi Zheng , Zhen Sun , Yuemeng Zhao , Yule Liu , Xinlei He , Xinyi Huang

This study explores the potential of Large Language Models (LLMs), specifically GPT-4, to enhance objectivity in organizational task performance evaluations. Through comparative analyses across two studies, including various task…

计算与语言 · 计算机科学 2024-08-13 Ning Li , Huaikang Zhou , Mingze Xu

In this work, we conduct an assessment of the optimization capabilities of LLMs across various tasks and data sizes. Each of these tasks corresponds to unique optimization domains, and LLMs are required to execute these tasks with…

机器学习 · 计算机科学 2024-05-28 Pei-Fu Guo , Ying-Hsuan Chen , Yun-Da Tsai , Shou-De Lin

Ideal or real - that is the question.In this work, we explore whether principles from game theory can be effectively applied to the evaluation of large language models (LLMs). This inquiry is motivated by the growing inadequacy of…

计算与语言 · 计算机科学 2026-04-07 Gao Yang , Yuhang Liu , Siyu Miao , Xinyue Liang , Zhengyang Liu , Heyan Huang

Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that differ from real-world usage in task-oriented dialogue (TOD). In this setting, LLMs must…

计算与语言 · 计算机科学 2026-04-30 Ivan Kartáč , Mateusz Lango , Ondřej Dušek

This paper explores the frontiers of large language models (LLMs) in psychology applications. Psychology has undergone several theoretical changes, and the current use of Artificial Intelligence (AI) and Machine Learning, particularly LLMs,…

机器学习 · 计算机科学 2025-07-15 Luoma Ke , Song Tong , Peng Cheng , Kaiping Peng
‹ 上一页 1 8 9 10 下一页 ›