English
Related papers

Related papers: Detecting Strategic Deception Using Linear Probes

200 papers

Can deception be detected solely from written text? Cues of deceptive communication are inherently subtle, even more so in text-only communication. Yet, prior studies have reported considerable success in automatic deception detection. We…

Computation and Language · Computer Science 2026-02-18 Aswathy Velutharambath , Kai Sassenberg , Roman Klinger

Reliable detection of deceptive behavior in Large Language Model (LLM) agents is an essential prerequisite for safe deployment in high-stakes agentic contexts. Prior work on scheming detection has focused exclusively on black-box monitors…

Computation and Language · Computer Science 2026-03-17 Snehasis Mukhopadhyay

LLMs often produce fluent but incorrect answers, yet detecting such hallucinations typically requires multiple sampling passes or post-hoc verification, adding significant latency and cost. We hypothesize that intermediate layers encode…

Computation and Language · Computer Science 2026-01-30 Rohan Bhatnagar , Youran Sun , Chi Andrew Zhang , Yixin Wen , Haizhao Yang

Large language models (LLMs) can provide users with false, inaccurate, or misleading information, and we consider the output of this type of information as what Natale (2021) calls `banal' deceptive behaviour. Here, we investigate peoples'…

Computers and Society · Computer Science 2025-10-29 Xiao Zhan , Yifan Xu , Noura Abdi , Joe Collenette , Ruba Abu-Salma , Stefan Sarkadi

Language Models (LMs) may acquire harmful knowledge, and yet feign ignorance of these topics when under audit. Inspired by the recent discovery of deception-related behaviour patterns in LMs, we aim to train classifiers that detect when a…

Computation and Language · Computer Science 2026-03-24 Dhananjay Ashok , Ruth-Ann Armstrong , Jonathan May

Social media platforms like Twitter, Facebook, and Instagram have facilitated the spread of misinformation, necessitating automated detection systems. This systematic review evaluates 36 studies that apply machine learning (ML) and deep…

Machine Learning · Computer Science 2025-06-24 Yunchong Liu , Xiaorui Shen , Yeyubei Zhang , Zhongyan Wang , Yexin Tian , Jianglai Dai , Yuchen Cao

The ability to discern between true and false information is essential to making sound decisions. However, with the recent increase in AI-based disinformation campaigns, it has become critical to understand the influence of deceptive…

Computers and Society · Computer Science 2022-10-18 Valdemar Danry , Pat Pataranutaporn , Ziv Epstein , Matthew Groh , Pattie Maes

As VLMs are deployed in safety-critical applications, their ability to abstain from answering when uncertain becomes crucial for reliability, especially in Scene Text Visual Question Answering (STVQA) tasks. For example, OCR errors like…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Jihan Yao , Achin Kulshrestha , Nathalie Rauschmayr , Reed Roberts , Banghua Zhu , Yulia Tsvetkov , Federico Tombari

Deception detection has attracted increasing attention due to its importance in real-world scenarios. Its main goal is to detect deceptive behaviors from multimodal clues such as gestures, facial expressions, prosody, etc. However, these…

Computation and Language · Computer Science 2024-08-14 Kang Chen , Zheng Lian , Haiyang Sun , Rui Liu , Jiangyan Yi , Bin Liu , Jianhua Tao

Large Language Models (LLMs) frequently exhibit unfaithful behavior, producing a final answer that differs significantly from their internal chain of thought (CoT) reasoning in order to appease the user they are conversing with. In order to…

Computation and Language · Computer Science 2026-02-04 Shikhar Shiromani , Archie Chaudhury , Sri Pranav Kunda

Large language models sometimes produce false or misleading responses. Two approaches to this problem are honesty elicitation -- modifying prompts or weights so that the model answers truthfully -- and lie detection -- classifying whether a…

Machine Learning · Computer Science 2026-03-11 Helena Casademunt , Bartosz Cywiński , Khoi Tran , Arya Jakkli , Samuel Marks , Neel Nanda

How to detect and mitigate deceptive AI systems is an open problem for the field of safe and trustworthy AI. We analyse two algorithms for mitigating deception: The first is based on the path-specific objectives framework where paths in the…

Artificial Intelligence · Computer Science 2023-06-27 Ismail Sahbane , Francis Rhys Ward , C Henrik Åslund

Most commonly used language models (LMs) are instruction-tuned and aligned using a combination of fine-tuning and reinforcement learning, causing them to refuse users requests deemed harmful by the model. However, jailbreak prompts can…

Computation and Language · Computer Science 2025-07-02 Aryan Shrivastava , Ari Holtzman

Lie detection is considered a concern for everyone in their day to day life given its impact on human interactions. Thus, people normally pay attention to both what their interlocutors are saying and also to their visual appearances,…

Computer Vision and Pattern Recognition · Computer Science 2021-07-01 Nuria Rodriguez-Diaz , Decky Aspandi , Federico Sukno , Xavier Binefa

Explanations for AI models in high-stakes domains like medicine often lack verifiability, which can hinder trust. To address this, we propose an interactive agent that produces explanations through an auditable sequence of actions. The…

Artificial Intelligence · Computer Science 2025-11-04 Yuhang Huang , Zekai Lin , Fan Zhong , Lei Liu

As Large Language Models (LLMs) become increasingly integrated into our daily lives, the potential harms from deceptive behavior underlie the need for faithfully interpreting their decision-making. While traditional probing methods have…

Machine Learning · Computer Science 2024-11-08 Anthony Costarelli , Mat Allen , Severin Field

Do large language models (LLMs) anticipate when they will answer correctly? To study this, we extract activations after a question is read but before any tokens are generated, and train linear probes to predict whether the model's…

Large Language Models (LLMs) have revolutionised natural language processing, exhibiting impressive human-like capabilities. In particular, LLMs are capable of "lying", knowingly outputting false statements. Hence, it is of interest and…

Computation and Language · Computer Science 2024-10-22 Lennart Bürger , Fred A. Hamprecht , Boaz Nadler

Recent works for time-series forecasting more and more leverage the high predictive power of Deep Learning models. With this increase in model complexity, however, comes a lack in understanding of the underlying model decision process,…

Machine Learning · Computer Science 2025-01-17 Matthias Jakobs , Thomas Liebig

Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typically validated in narrow settings that do not capture the…

Computation and Language · Computer Science 2026-01-12 Kieron Kretschmar , Walter Laurito , Sharan Maiya , Samuel Marks