English
Related papers

Related papers: Deceive, Detect, and Disclose: Large Language Mode…

200 papers

Large Language Model (LLM) agents are increasingly used in many applications, raising concerns about their safety. While previous work has shown that LLMs can deceive in controlled tasks, less is known about their ability to deceive using…

Artificial Intelligence · Computer Science 2026-01-21 Christopher Kao , Vanshika Vats , James Davis

In this paper, we study a game called ``Mafia,'' in which different players have different types of information, communication and functionality. The players communicate and function in a way that resembles some real-life situations. We…

Probability · Mathematics 2009-09-29 Mark Braverman , Omid Etesami , Elchanan Mossel

While neural networks demonstrate a remarkable ability to model linguistic content, capturing contextual information related to a speaker's conversational role is an open area of research. In this work, we analyze the effect of speaker role…

Computation and Language · Computer Science 2022-07-07 Samee Ibraheem , Gaoyue Zhou , John DeNero

Social deduction games such as Mafia present a unique AI challenge: players must reason under uncertainty, interpret incomplete and intentionally misleading information, evaluate human-like communication, and make strategic elimination…

Artificial Intelligence · Computer Science 2026-04-22 Mihir Shriniwas Arya , Avinash Anish , Aditya Ranjan

Are current language models capable of deception and lie detection? We study this question by introducing a text-based game called $\textit{Hoodwinked}$, inspired by Mafia and Among Us. Players are locked in a house and must find a key to…

Computation and Language · Computer Science 2023-08-07 Aidan O'Gara

Quantifying the deceptive potential of Large Language Models (LLMs) is critical for AI safety, yet difficult to achieve in uncontrolled environments. This work investigates the reasoning, persuasion, and deceptive capabilities of LLMs…

Computation and Language · Computer Science 2026-05-25 Niklas Bauer

The behavior of Large Language Models (LLMs) as artificial social agents is largely unexplored, and we still lack extensive evidence of how these agents react to simple social stimuli. Testing the behavior of AI agents in classic Game…

Computers and Society · Computer Science 2024-09-20 Nicoló Fontana , Francesco Pierri , Luca Maria Aiello

Detecting deception in natural language has a wide variety of applications, but because of its hidden nature there are currently no public, large-scale sources of labeled deceptive text. This work introduces the Mafiascum dataset [1], a…

Computation and Language · Computer Science 2019-08-15 Bob de Ruiter , George Kachergis

As language models are deployed as autonomous agents that negotiate, cooperate, and compete on behalf of human principals, their strategic dispositions acquire direct economic consequences. Here we show, across 51,906 game-theoretic trials…

Physics and Society · Physics 2026-04-23 Felipe M. Affonso

In Social Deduction Games (SDGs) such as Avalon, Mafia, and Werewolf, players conceal their identities and deliberately mislead others, making hidden-role inference a central and demanding task. Accurate role identification, which forms the…

Artificial Intelligence · Computer Science 2025-11-11 Kaijie Xu , Fandi Meng , Clark Verbrugge , Simon Lucas

Large language model-based (LLM-based) agents have become common in settings that include non-cooperative parties. In such settings, agents' decision-making needs to conceal information from their adversaries, reveal information to their…

Artificial Intelligence · Computer Science 2025-10-22 Mustafa O. Karabag , Jan Sobotka , Ufuk Topcu

As Large Language Models (LLMs) transition into autonomous agentic roles, the risk of deception-defined behaviorally as the systematic provision of false information to satisfy external incentives-poses a significant challenge to AI safety.…

Computation and Language · Computer Science 2026-03-10 Arash Marioriyad , Ali Nouri , Mohammad Hossein Rohban , Mahdieh Soleymani Baghshah

Deception is a fundamental challenge for multi-agent reasoning: effective systems must strategically conceal information while detecting misleading behavior in others. Yet most evaluations reduce deception to static classification, ignoring…

Multiagent Systems · Computer Science 2025-12-11 Mrinal Agarwal , Saad Rana , Theo Sundoro , Hermela Berhe , Spencer Kim , Vasu Sharma , Sean O'Brien , Kevin Zhu

Large language models are increasingly deployed as autonomous agents in multi-agent settings where they communicate intentions and take consequential actions with limited human oversight. A critical safety question is whether agents that…

Computers and Society · Computer Science 2026-04-07 Jerick Shi , Terry Jingcheng Zhang , Zhijing Jin , Vincent Conitzer

This study reveals how frontier Large Language Models LLMs can "game the system" when faced with impossible situations, a critical security and alignment concern. Using a novel textual simulation approach, we presented three leading LLMs…

Artificial Intelligence · Computer Science 2025-05-14 Lars Malmqvist

Social reasoning - inferring unobservable beliefs and intentions from partial observations of other agents - remains a challenging task for large language models (LLMs). We evaluate the limits of current reasoning language models in the…

Artificial Intelligence · Computer Science 2026-04-13 Shahab Rahimirad , Guven Gergerli , Lucia Romero , Angela Qian , Matthew Lyle Olson , Simon Stepputtis , Joseph Campbell

We demonstrate LLM agent specification gaming by instructing models to win against a chess engine. We find reasoning models like OpenAI o3 and DeepSeek R1 will often hack the benchmark by default, while language models like GPT-4o and…

Artificial Intelligence · Computer Science 2025-08-28 Alexander Bondarenko , Denis Volk , Dmitrii Volkov , Jeffrey Ladish

Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire increased agency and human oversight diminishes. In this work,…

Artificial Intelligence · Computer Science 2026-03-10 Matthew Lyle Olson , Neale Ratzlaff , Musashi Hinck , Tri Nguyen , Vasudev Lal , Joseph Campbell , Simon Stepputtis , Shao-Yen Tseng

Mafia can be described as an experiment in human psychology and mass hysteria, or as a game between informed minority and uninformed majority. Focus on a very restricted setting, Mossel et al. [to appear in Ann. Appl. Probab. Volume 18,…

Probability · Mathematics 2008-04-02 Erlin Yao

The rapid advancement of Large Language Models (LLMs) has necessitated more robust evaluation methods that go beyond static benchmarks, which are increasingly prone to data saturation and leakage. In this paper, we propose a dynamic…

Computation and Language · Computer Science 2026-01-15 Haryo Akbarianto Wibowo , Alaa Elsetohy , Qinrong Cui , Alham Fikri Aji
‹ Prev 1 2 3 10 Next ›