English
Related papers

Related papers: Information-theoretic Distinctions Between Decepti…

200 papers

The advances in artificial intelligence enabled by deep learning architectures are undeniable. In several cases, deep neural network driven models have surpassed human level performance in benchmark autonomy tasks. The underlying policies…

Artificial Intelligence · Computer Science 2021-06-11 Jeff Druce , James Niehaus , Vanessa Moody , David Jensen , Michael L. Littman

Large Language Models (LLMs) act as powerful reasoning engines but struggle with "symbol grounding" in embodied environments, particularly when information is asymmetrically distributed. We investigate the Privileged Information Bias (or…

Artificial Intelligence · Computer Science 2025-12-19 Shaun Baek , Sam Liu , Joseph Ukpong

Large language models (LLMs) are increasingly tasked with strategic decision-making under incomplete information, such as in negotiation and policymaking. While LLMs can excel at many such tasks, they also fail in ways that are poorly…

Computation and Language · Computer Science 2026-05-04 Jan Sobotka , Mustafa O. Karabag , Ufuk Topcu

AI models are already deployed in societies affected by armed conflict, and journalists, humanitarian workers, governments and ordinary citizens rely on them for information or for their work processes. No established practice exists for…

Artificial Intelligence · Computer Science 2026-05-22 Andrii Kryshtal

As AI systems increasingly assume roles where trust and alignment with human values are essential, understanding when and why they engage in deception has become a critical research priority. We introduce The Traitors, a multi-agent…

Artificial Intelligence · Computer Science 2025-12-16 Pedro M. P. Curvo

Current AI systems, grounded in oversimplified neuroscience, risk eroding the distinction between truth and falsehood. They maximize reward by amplifying attention to information without intrinsic precision mechanisms to assess whether it…

Artificial Intelligence · Computer Science 2026-05-05 Ahsan Adeel

We introduce the Adversarial Confusion Attack, a new class of threats against multimodal large language models (MLLMs). Unlike jailbreaks or targeted misclassification, the goal is to induce systematic disruption that makes the model…

Computation and Language · Computer Science 2025-12-02 Jakub Hoscilowicz , Artur Janicki

We study the design of autonomous agents that are capable of deceiving outside observers about their intentions while carrying out tasks in stochastic, complex environments. By modeling the agent's behavior as a Markov decision process, we…

Artificial Intelligence · Computer Science 2021-09-15 Yagiz Savas , Christos K. Verginis , Ufuk Topcu

As AI systems become increasingly capable and influential, ensuring their alignment with human values, preferences, and goals has become a critical research focus. Current alignment methods primarily focus on designing algorithms and loss…

Computation and Language · Computer Science 2025-05-02 Min-Hsuan Yeh , Jeffrey Wang , Xuefeng Du , Seongheon Park , Leitian Tao , Shawn Im , Yixuan Li

Large language models (LLMs) can provide users with false, inaccurate, or misleading information, and we consider the output of this type of information as what Natale (2021) calls `banal' deceptive behaviour. Here, we investigate peoples'…

Computers and Society · Computer Science 2025-10-29 Xiao Zhan , Yifan Xu , Noura Abdi , Joe Collenette , Ruba Abu-Salma , Stefan Sarkadi

Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here…

Physics and Society · Physics 2026-05-12 Giordano De Marzo , Alessandro Bellina , Claudio Castellano , Viola Priesemann , David Garcia

Can deception exist in differential games? We provide a case study for a Turret-Attacker differential game, where two Attackers seek to score points by reaching a target region while a Turret tries to minimize the score by aligning itself…

Computer Science and Game Theory · Computer Science 2024-05-14 Daigo Shishika , Alexander Von Moll , Dipankar Maity , Michael Dorothy

Meta-learning, or "learning to learn", refers to techniques that infer an inductive bias from data corresponding to multiple related tasks with the goal of improving the sample efficiency for new, previously unobserved, tasks. A key…

Machine Learning · Computer Science 2021-02-24 Sharu Theresa Jose , Osvaldo Simeone

Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding misinformation and persuasion do not adequately address. This paper proposes that a significant…

Human-Computer Interaction · Computer Science 2026-05-27 Andrew D. Maynard

Language Confusion is a phenomenon where Large Language Models (LLMs) generate text that is neither in the desired language, nor in a contextually appropriate language. This phenomenon presents a critical challenge in text generation by…

Computation and Language · Computer Science 2025-02-11 Yiyi Chen , Qiongxiu Li , Russa Biswas , Johannes Bjerva

Establishing shared goals is a fundamental step in human-AI communication. However, ambiguities can lead to outputs that seem correct but fail to reflect the speaker's intent. In this paper, we explore this issue with a focus on the data…

Computation and Language · Computer Science 2025-10-13 Mert İnan , Anthony Sicilia , Alex Xie , Saujas Vaduguru , Daniel Fried , Malihe Alikhani

Artificial intelligence (AI) is increasingly integrated into modern healthcare, offering powerful support for clinical decision-making. However, in real-world settings, AI systems may experience performance degradation over time, due to…

Artificial Intelligence · Computer Science 2026-02-05 Hao Guan , David Bates , Li Zhou

There is much discussion of the false outputs that generative AI systems such as ChatGPT, Claude, Gemini, DeepSeek, and Grok create. In popular terminology, these have been dubbed AI hallucinations. However, deeming these AI outputs…

Computers and Society · Computer Science 2026-04-17 Lucy Osler

Information-theoretic (IT) measures are ubiquitous in artificial intelligence: entropy drives decision-tree splits and uncertainty quantification, cross-entropy is the default classification loss, mutual information underpins representation…

Artificial Intelligence · Computer Science 2026-04-28 Nikolaos Al. Papadopoulos , Konstantinos E. Psannis

Goal-oriented conversational systems require making sequential decisions under uncertainty about the user's intent, where the algorithm must balance information acquisition and target commitment over multiple turns. Existing approaches…

Computation and Language · Computer Science 2026-04-07 Xinyi Ling , Ye Liu , Reza Averly , Xia Ning