English
Related papers

Related papers: Detecting Strategic Deception Using Linear Probes

200 papers

Despite the remarkable advances of Large Language Models (LLMs) across diverse cognitive tasks, the rapid enhancement of these capabilities also introduces emergent deceptive behaviors that may induce severe risks in high-stakes…

Computation and Language · Computer Science 2025-11-18 Yao Huang , Yitong Sun , Yichi Zhang , Ruochen Zhang , Yinpeng Dong , Xingxing Wei

Lies and deception are common phenomena in society, both in our private and professional lives. However, humans are notoriously bad at accurate deception detection. Based on the literature, human accuracy of distinguishing between lies and…

Computer Vision and Pattern Recognition · Computer Science 2018-12-31 Minh Ngô , Burak Mandira , Selim Fırat Yılmaz , Ward Heij , Sezer Karaoglu , Henri Bouma , Hamdi Dibeklioglu , Theo Gevers

As AI systems increasingly assume roles where trust and alignment with human values are essential, understanding when and why they engage in deception has become a critical research priority. We introduce The Traitors, a multi-agent…

Artificial Intelligence · Computer Science 2025-12-16 Pedro M. P. Curvo

Large language models are being widely used across industries to generate content that contributes directly to key performance metrics, such as conversion rates. Pretrained models, however, often fall short when it comes to aligning with…

Machine Learning · Computer Science 2025-06-03 Erfan Loghmani

As large language models (LLMs) are increasingly deployed as autonomous agents, understanding how strategic behavior emerges in multi-agent environments has become an important alignment challenge. We take a neutral empirical stance and…

The detection of political fake statements is crucial for maintaining information integrity and preventing the spread of misinformation in society. Historically, state-of-the-art machine learning models employed various methods for…

Computation and Language · Computer Science 2023-06-16 Mars Gokturk Buchholz

While Large Language Models (LLMs) have shown exceptional performance in various tasks, one of their most prominent drawbacks is generating inaccurate or false information with a confident tone. In this paper, we provide evidence that the…

Computation and Language · Computer Science 2023-10-18 Amos Azaria , Tom Mitchell

It is difficult for humans to distinguish the true and false of rumors, but current deep learning models can surpass humans and achieve excellent accuracy on many rumor datasets. In this paper, we investigate whether deep learning models…

Computation and Language · Computer Science 2022-05-31 Shiwen Ni , Jiawen Li , Hung-Yu Kao

Large Language Models (LLMs) have been demonstrating strong reasoning capability with their chain-of-thoughts (CoT), which are routinely used by humans to judge answer quality. This reliance creates a powerful yet fragile basis for trust.…

Machine Learning · Computer Science 2026-05-22 Wei Shen , Han Wang , Haoyu Li , Huan Zhang

Despite the importance of developing generative AI models that can effectively resist scams, current literature lacks a structured framework for evaluating their vulnerability to such threats. In this work, we address this gap by…

Cryptography and Security · Computer Science 2025-07-18 Udari Madhushani Sehwag , Kelly Patel , Francesca Mosca , Vineeth Ravi , Jessica Staddon

System prompts for AI coding agents increasingly employ motivational framing -- from neutral task descriptions to fear-driven threats -- yet no controlled study has examined whether such framing affects agent behavior. We present two…

Software Engineering · Computer Science 2026-03-17 Wu Ji

Among areas of software engineering where AI techniques -- particularly, Large Language Models -- seem poised to yield dramatic improvements, an attractive candidate is Automatic Program Repair (APR), the production of satisfactory…

Software Engineering · Computer Science 2025-08-05 Li Huang , Ilgiz Mustafin , Marco Piccioni , Alessandro Schena , Reto Weber , Bertrand Meyer

Alignment audits aim to robustly identify hidden goals from strategic, situationally aware misaligned models. Despite this threat model, existing auditing methods have not been systematically stress-tested against deception strategies. We…

Machine Learning · Computer Science 2026-03-09 Oliver Daniels , Perusha Moodley , Benjamin M. Marlin , David Lindner

Decentralized AI agent networks, such as Gaia, allows individuals to run customized LLMs on their own computers and then provide services to the public. However, in order to maintain service quality, the network must verify that individual…

Artificial Intelligence · Computer Science 2025-08-20 Michael J. Yuan , Carlos Lospoy , Sydney Lai , James Snewin , Ju Long

Probing has emerged as a promising method for monitoring large language models (LLMs), enabling cheap inference-time detection of concerning behaviours. However, natural examples of many behaviours are rare, forcing researchers to rely on…

Artificial Intelligence · Computer Science 2026-04-21 Nathalie Kirch , Samuel Dower , Adrians Skapars , Helen Yannakoudakis , Ekdeep Singh Lubana , Dmitrii Krasheninnikov

In order for AI systems to communicate effectively with people, they must understand how we make decisions. However, people's decisions are not always rational, so the implicit internal models of human decision-making in Large Language…

Computation and Language · Computer Science 2025-03-11 Ryan Liu , Jiayi Geng , Joshua C. Peterson , Ilia Sucholutsky , Thomas L. Griffiths

In safety-critical applications, language models should be able to characterize their uncertainty with meaningful probabilities. Many uncertainty quantification approaches require supervised data; however, finding suitable unseen…

Computation and Language · Computer Science 2026-05-14 Sophia Hager , Simon Zeng , Nicholas Andrews

Monitoring autonomous large language model (LLM) agents for covert malicious behavior is challenging due to delayed, context-dependent, and long-horizon attack patterns. Agents may pursue hidden objectives while maintaining superficially…

Machine Learning · Computer Science 2026-05-26 Nesreen K. Ahmed , Nima Nafisi

Large Language Models are increasingly proposed as cognitive components for robotic systems, yet their opaque decision processes make it difficult to explain success or failure in closed-loop embodied tasks. Following an empirical AI…

Artificial Intelligence · Computer Science 2026-05-20 Oussama Zenkri , Oliver Brock

The safety and alignment of Large Language Models (LLMs) are critical for their responsible deployment. Current evaluation methods predominantly focus on identifying and preventing overtly harmful outputs. However, they often fail to…