English
Related papers

Related papers: Among Us: A Sandbox for Measuring and Detecting Ag…

200 papers

Despite the remarkable advances of Large Language Models (LLMs) across diverse cognitive tasks, the rapid enhancement of these capabilities also introduces emergent deceptive behaviors that may induce severe risks in high-stakes…

Computation and Language · Computer Science 2025-11-18 Yao Huang , Yitong Sun , Yichi Zhang , Ruochen Zhang , Yinpeng Dong , Xingxing Wei

Deception and persuasion play a critical role in long-horizon dialogues between multiple parties, especially when the interests, goals, and motivations of the participants are not aligned. Such complex tasks pose challenges for current…

Computation and Language · Computer Science 2023-11-13 Simon Stepputtis , Joseph Campbell , Yaqi Xie , Zhengyang Qi , Wenxin Sharon Zhang , Ruiyi Wang , Sanketh Rangreji , Michael Lewis , Katia Sycara

Simulating high quality user behavior data has always been a fundamental problem in human-centered applications, where the major difficulty originates from the intricate mechanism of human decision process. Recently, substantial evidences…

Information Retrieval · Computer Science 2024-02-16 Lei Wang , Jingsen Zhang , Hao Yang , Zhiyuan Chen , Jiakai Tang , Zeyu Zhang , Xu Chen , Yankai Lin , Ruihua Song , Wayne Xin Zhao , Jun Xu , Zhicheng Dou , Jun Wang , Ji-Rong Wen

Deceptive agents are a challenge for the safety, trustworthiness, and cooperation of AI systems. We focus on the problem that agents might deceive in order to achieve their goals (for instance, in our experiments with language models, the…

Artificial Intelligence · Computer Science 2023-12-05 Francis Rhys Ward , Francesco Belardinelli , Francesca Toni , Tom Everitt

In this paper, we explore the potential of Large Language Models (LLMs) Agents in playing the strategic social deduction game, Resistance Avalon. Players in Avalon are challenged not only to make informed decisions based on dynamically…

Artificial Intelligence · Computer Science 2023-11-09 Jonathan Light , Min Cai , Sheng Shen , Ziniu Hu

As large language models (LLMs) become more capable and agentic, the requirement for trust in their outputs grows significantly, yet at the same time concerns have been mounting that models may learn to lie in pursuit of their goals. To…

With ChatGPT-like large language models (LLM) prevailing in the community, how to evaluate the ability of LLMs is an open question. Existing evaluation methods suffer from following shortcomings: (1) constrained evaluation abilities, (2)…

Artificial Intelligence · Computer Science 2023-08-09 Jiaju Lin , Haoran Zhao , Aochi Zhang , Yiting Wu , Huqiuyue Ping , Qin Chen

The rapid advancement of Large Language Models (LLMs) has necessitated more robust evaluation methods that go beyond static benchmarks, which are increasingly prone to data saturation and leakage. In this paper, we propose a dynamic…

Computation and Language · Computer Science 2026-01-15 Haryo Akbarianto Wibowo , Alaa Elsetohy , Qinrong Cui , Alham Fikri Aji

Recent advances in Large Language Models (LLMs) have incorporated planning and reasoning capabilities, enabling models to outline steps before execution and provide transparent reasoning paths. This enhancement has reduced errors in…

Computation and Language · Computer Science 2025-01-31 Sudarshan Kamath Barkur , Sigurd Schacht , Johannes Scholl

Large language models (LLMs) can "lie", which we define as outputting false statements despite "knowing" the truth in a demonstrable sense. LLMs might "lie", for example, when instructed to output misinformation. Here, we develop a simple…

Computation and Language · Computer Science 2023-09-28 Lorenzo Pacchiardi , Alex J. Chan , Sören Mindermann , Ilan Moscovitz , Alexa Y. Pan , Yarin Gal , Owain Evans , Jan Brauner

The potential of Large Language Model (LLM) as agents has been widely acknowledged recently. Thus, there is an urgent need to quantitatively \textit{evaluate LLMs as agents} on challenging tasks in interactive environments. We present…

Humans are capable of strategically deceptive behavior: behaving helpfully in most situations, but then behaving very differently in order to pursue alternative objectives when given the opportunity. If an AI system learned such a deceptive…

Large Language Models (LLMs) have demonstrated an alarming ability to impersonate humans in conversation, raising concerns about their potential misuse in scams and deception. Humans have a right to know if they are conversing to an LLM. We…

Computation and Language · Computer Science 2024-12-23 Gilad Gressel , Rahul Pankajakshan , Yisroel Mirsky

Large Language Model (LLM)-based agents are increasingly used as autonomous subordinates that carry out tasks for users. This raises the question of whether they may also engage in deception, similar to how individuals in human…

The emergence of autonomous Large Language Model (LLM) agents capable of tool usage has introduced new safety risks that go beyond traditional conversational misuse. These agents, empowered to execute external functions, are vulnerable to…

Artificial Intelligence · Computer Science 2025-07-14 Zeyang Sha , Hanling Tian , Zhuoer Xu , Shiwen Cui , Changhua Meng , Weiqiang Wang

As large language model agents advance beyond software engineering (SWE) tasks toward machine learning engineering (MLE), verifying agent behavior becomes orders of magnitude more expensive: while SWE tasks can be verified via…

Computation and Language · Computer Science 2026-04-07 Yuhang Zhou , Lizhu Zhang , Yifan Wu , Jiayi Liu , Xiangjun Fan , Zhuokai Zhao , Hong Yan

Evaluating the safety of frontier AI systems is an increasingly important concern, helping to measure the capabilities of such models and identify risks before deployment. However, it has been recognised that if AI agents are aware that…

Machine Learning · Computer Science 2025-10-01 Joel Dyer , Daniel Jarne Ornia , Nicholas Bishop , Anisoara Calinescu , Michael Wooldridge

Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Model (LLM) agents. However, most environments are scored only by game outcomes such as win rates and…

Computation and Language · Computer Science 2026-05-27 Ye Yuan , Rui Song , Weien Li , Zeyu Li , Haochen Liu , Xiangyu Kong , Changjiang Han , Yonghan Yang , Zichen Zhao , Zixuan Dong , Fuyuan Lyu , Bowei He , Haolun Wu , Jikun Kang , Xue Liu

Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments. This work investigates the reasoning, persuasion, and deceptive capabilities of LLMs…

Computation and Language · Computer Science 2026-05-25 Niklas Bauer