English
Related papers

Related papers: The Decrypto Benchmark for Multi-Agent Reasoning a…

200 papers

Reasoning is a fundamental cognitive process that enables logical inference, problem-solving, and decision-making. With the rapid advancement of large language models (LLMs), reasoning has emerged as a key capability that distinguishes…

Recent advancements in large language models (LLMs) have revealed their potential for achieving autonomous agents possessing human-level intelligence. However, existing benchmarks for evaluating LLM Agents either use static datasets,…

Computation and Language · Computer Science 2024-02-27 Junzhe Chen , Xuming Hu , Shuodi Liu , Shiyu Huang , Wei-Wei Tu , Zhaofeng He , Lijie Wen

Theory of Mind (ToM), the ability to understand the mental states of oneself and others, remains a challenging area for large language models (LLMs), which often fail to predict human mental states accurately. In this paper, we introduce…

Computation and Language · Computer Science 2025-06-12 Prameshwar Thiyagarajan , Vaishnavi Parimi , Shamant Sai , Soumil Garg , Zhangir Meirbek , Nitin Yarlagadda , Kevin Zhu , Chris Kim

Theory of Mind (ToM), the ability to track others epistemic state, makes humans efficient collaborators. AI agents need the same capacity in multi agent settings, yet existing benchmarks mostly test literal ToM by asking direct belief…

Artificial Intelligence · Computer Science 2026-05-19 Gurusha Juneja , Dylan Lu , Saaket Agashe , Parth Diwane , Edward Gunn , Jayanth Srinivasa , Gaowen Liu , William Yang Wang , Yali Du , Xin Eric Wang

Do large language models (LLMs) have theory of mind? A plethora of papers and benchmarks have been introduced to evaluate if current models have been able to develop this key ability of social intelligence. However, all rely on limited…

Machine Learning · Computer Science 2024-12-18 Melanie Sclar , Jane Yu , Maryam Fazel-Zarandi , Yulia Tsvetkov , Yonatan Bisk , Yejin Choi , Asli Celikyilmaz

Large Language models are revolutionizing the conversational recommender systems through their impressive capabilities in instruction comprehension, reasoning, and human interaction. A core factor underlying effective recommendation…

Artificial Intelligence · Computer Science 2025-12-01 Mengfan Li , Xuanhua Shi , Yang Deng

Hallucination continues to pose a major obstacle in the reasoning capabilities of large language models (LLMs). Although the Multi-Agent Debate (MAD) paradigm offers a promising solution by promoting consensus among multiple agents to…

Artificial Intelligence · Computer Science 2025-11-17 Dayong Liang , Xiao-Yong Wei , Changmeng Zheng

The rapid advancement of large language models (LLMs) has enabled role-playing language agents to demonstrate significant potential in various applications. However, relying solely on prompts and contextual inputs often proves insufficient…

Computation and Language · Computer Science 2025-07-24 Xiaoyu Zhan , Xinyu Fu , Hao Sun , Yuanqi Li , Jie Guo , Yanwen Guo

Can emergent language models faithfully model the intelligence of decision-making agents? Though modern language models exhibit already some reasoning ability, and theoretically can potentially express any probable distribution over tokens,…

Machine Learning · Computer Science 2024-06-27 Wenhao Lu , Xufeng Zhao , Josua Spisak , Jae Hee Lee , Stefan Wermter

Theory of Mind (ToM) refers to the ability to infer others' mental states, such as beliefs, desires, and intentions. Current vision-language embodied agents lack ToM-based decision-making, and existing benchmarks focus solely on human…

Artificial Intelligence · Computer Science 2026-02-25 Ruoxuan Zhang , Qiyun Zheng , Zhiyu Zhou , Ziqi Liao , Siyu Wu , Jian-Yu Jiang-Lin , Bin Wen , Hongxia Xie , Jianlong Fu , Wen-Huang Cheng

We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning--the proactive construction, testing, and revision of hypotheses in information-sparse environments. Existing benchmarks, often static or focused on…

Artificial Intelligence · Computer Science 2025-08-15 Mengtao Zhou , Sifan Wu , Huan Zhang , Qi Sima , Bang Liu

Large language models (LLMs) can exhibit biases in reasoning capabilities due to linguistic modality, performing better on tasks in one language versus another, even with similar content. Most previous works evaluate this through reasoning…

Computation and Language · Computer Science 2025-10-17 César Guerra-Solano , Zhuochun Li , Xiang Lorraine Li

Large language models (LLMs) are increasingly used as judges to evaluate agent performance, particularly in non-verifiable settings where judgments rely on agent trajectories including chain-of-thought (CoT) reasoning. This paradigm…

Artificial Intelligence · Computer Science 2026-01-23 Muhammad Khalifa , Lajanugen Logeswaran , Jaekyeom Kim , Sungryull Sohn , Yunxiang Zhang , Moontae Lee , Hao Peng , Lu Wang , Honglak Lee

We introduce LLM CHESS, an evaluation framework designed to probe the generalization of reasoning and instruction-following abilities in large language models (LLMs) through extended agentic interaction in the domain of chess. We rank over…

Artificial Intelligence · Computer Science 2025-12-02 Sai Kolasani , Maxim Saplin , Nicholas Crispino , Kyle Montgomery , Jared Quincy Davis , Matei Zaharia , Chi Wang , Chenguang Wang

The question of whether large language models (LLMs) possess Theory of Mind (ToM) -- often defined as the ability to reason about others' mental states -- has sparked significant scientific and public interest. However, the evidence as to…

Artificial Intelligence · Computer Science 2025-03-03 Jennifer Hu , Felix Sosa , Tomer Ullman

Deducing whodunit proves challenging for LLM agents. In this paper, we implement a text-based multi-agent version of the classic board game Clue as a rule-based testbed for evaluating multi-step deductive reasoning, with six agents drawn…

Artificial Intelligence · Computer Science 2026-03-19 Rebecca Ansell , Autumn Toney-Wails

As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important. To fully assess LLM intelligence, it is crucial to examine their interactive dynamics…

Computation and Language · Computer Science 2025-09-23 Junhao Chen , Jingbo Sun , Xiang Li , Haidong Xin , Yuhao Xue , Yibin Xu , Hao Zhao

Large Language Models have shown exceptional generative abilities in various natural language and generation tasks. However, possible anthropomorphization and leniency towards failure cases have propelled discussions on emergent abilities…

Robotics · Computer Science 2024-01-18 Mudit Verma , Siddhant Bhambri , Subbarao Kambhampati

Understanding and reasoning over tables is a critical capability for many real-world applications. Large language models (LLMs) have shown promise on this task, but current approaches remain limited. Fine-tuning based methods strengthen…

‹ Prev 1 4 5 6 7 8 10 Next ›