English
Related papers

Related papers: Uncovering Deceptive Tendencies in Language Models…

200 papers

As ongoing research explores the ability of AI agents to be insider threats and act against company interests, we showcase the abilities of such agents to act against human well being in service of corporate authority. Building on Agentic…

Artificial Intelligence · Computer Science 2026-04-10 Thomas Rivasseau

Artificial intelligence (AI) comes with great opportunities but can also pose significant risks. Automatically generated explanations for decisions can increase transparency and foster trust, especially for systems based on automated…

Machine Learning · Computer Science 2021-12-03 Johannes Schneider , Christian Meske , Michalis Vlachos

We demonstrate how AI agents can coordinate to deceive oversight systems using automated interpretability of neural networks. Using sparse autoencoders (SAEs) as our experimental framework, we show that language models (Llama, DeepSeek R1,…

Artificial Intelligence · Computer Science 2025-04-11 Simon Lermen , Mateusz Dziemian , Natalia Pérez-Campanero Antolín

Large language models now possess human-level linguistic abilities in many contexts. This raises the concern that they can be used to deceive and manipulate on unprecedented scales, for instance spreading political misinformation on social…

Computers and Society · Computer Science 2026-01-21 Christian Tarsney

Large Language Models have become an integral part of new intelligent and interactive writing assistants. Many are offered commercially with a chatbot-like UI, such as ChatGPT, and provide little information about their inner workings. This…

Human-Computer Interaction · Computer Science 2024-04-16 Karim Benharrak , Tim Zindulka , Daniel Buschek

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system…

Large Language Models (LLMs) are able to provide assistance on a wide range of information-seeking tasks. However, model outputs may be misleading, whether unintentionally or in cases of intentional deception. We investigate the ability of…

Computation and Language · Computer Science 2024-07-17 Betty Li Hou , Kejian Shi , Jason Phang , James Aung , Steven Adler , Rosie Campbell

Humans and machines interact more frequently than ever and our societies are becoming increasingly hybrid. A consequence of this hybridisation is the degradation of societal trust due to the prevalence of AI-enabled deception. Yet, despite…

Multiagent Systems · Computer Science 2024-06-12 Stefan Sarkadi

Large Language Model (LLM)-based agents are increasingly used as autonomous subordinates that carry out tasks for users. This raises the question of whether they may also engage in deception, similar to how individuals in human…

The ability to discern between true and false information is essential to making sound decisions. However, with the recent increase in AI-based disinformation campaigns, it has become critical to understand the influence of deceptive…

Computers and Society · Computer Science 2022-10-18 Valdemar Danry , Pat Pataranutaporn , Ziv Epstein , Matthew Groh , Pattie Maes

Advanced Artificial Intelligence (AI) systems, specifically large language models (LLMs), have the capability to generate not just misinformation, but also deceptive explanations that can justify and propagate false information and erode…

Artificial Intelligence · Computer Science 2024-08-02 Valdemar Danry , Pat Pataranutaporn , Matthew Groh , Ziv Epstein , Pattie Maes

As large language models (LLMs) are increasingly deployed as interactive agents, open-ended human-AI interactions can involve deceptive behaviors with serious real-world consequences, yet existing evaluations remain largely…

Artificial Intelligence · Computer Science 2026-02-09 Yichen Wu , Qianqian Gao , Xudong Pan , Geng Hong , Min Yang

This paper argues that a range of current AI systems have learned how to deceive humans. We define deception as the systematic inducement of false beliefs in the pursuit of some outcome other than the truth. We first survey empirical…

Computers and Society · Computer Science 2023-08-29 Peter S. Park , Simon Goldstein , Aidan O'Gara , Michael Chen , Dan Hendrycks

Large Language Models (LLMs) can generate content that is as persuasive as human-written text and appear capable of selectively producing deceptive outputs. These capabilities raise concerns about potential misuse and unintended…

Computation and Language · Computer Science 2024-12-24 Cameron R. Jones , Benjamin K. Bergen

We demonstrate a situation in which Large Language Models, trained to be helpful, harmless, and honest, can display misaligned behavior and strategically deceive their users about this behavior without being instructed to do so. Concretely,…

Computation and Language · Computer Science 2024-07-16 Jérémy Scheurer , Mikita Balesni , Marius Hobbhahn

Recent research on large language models (LLMs) has demonstrated their ability to understand and employ deceptive behavior, even without explicit prompting. However, such behavior has only been observed in rare, specialized cases and has…

Computation and Language · Computer Science 2025-06-24 Laurène Vaugrante , Francesca Carlon , Maluna Menke , Thilo Hagendorff

Conversational Artificial Intelligence (AI) used in industry settings can be trained to closely mimic human behaviors, including lying and deception. However, lying is often a necessary part of negotiation. To address this, we develop a…

Computers and Society · Computer Science 2021-03-16 Tae Wan Kim , Tong , Lu , Kyusong Lee , Zhaoqi Cheng , Yanhan Tang , John Hooker

Text-based misinformation permeates online discourses, yet evidence of people's ability to discern truth from such deceptive textual content is scarce. We analyze a novel TV game show data where conversations in a high-stake environment…

Computation and Language · Computer Science 2024-04-09 Sanchaita Hazra , Bodhisattwa Prasad Majumder

Building reliable deception detectors for AI systems -- methods that could predict when an AI system is being strategically deceptive without necessarily requiring behavioural evidence -- would be valuable in mitigating risks from advanced…

Machine Learning · Computer Science 2025-12-17 Lewis Smith , Bilal Chughtai , Neel Nanda

Automated verbal deception detection using methods from Artificial Intelligence (AI) has been shown to outperform humans in disentangling lies from truths. Research suggests that transparency and interpretability of computational methods…

Human-Computer Interaction · Computer Science 2026-04-10 Riccardo Loconte , Merylin Monaro , Pietro Pietrini , Bruno Verschuere , Bennett Kleinberg
‹ Prev 1 2 3 10 Next ›