English
Related papers

Related papers: Detecting Strategic Deception Using Linear Probes

200 papers

Systems operating in adversarial environments may inadvertently leak sensitive information to adversaries. To address this challenge, we revisit the linear-quadratic control framework and introduce deception to actively mislead adversaries.…

Optimization and Control · Mathematics 2026-04-02 Yerin Kim , Haosheng Zhou , Alexander Benvenuti , Ruimeng Hu , Matthew Hale

Existing deep active learning algorithms achieve impressive sampling efficiency on natural language processing tasks. However, they exhibit several weaknesses in practice, including (a) inability to use uncertainty sampling with black-box…

Computation and Language · Computer Science 2020-07-22 Haw-Shiuan Chang , Shankar Vembu , Sunil Mohan , Rheeya Uppaal , Andrew McCallum

Millions of users turn to AI models for their information needs. It is conceivable that a large number of user queries contain assumptions that may be factually inaccurate. Prior work notes that large language models (LLMs) often fail to…

Computation and Language · Computer Science 2026-05-06 Rose Sathyanathan , Kinshuk Vasisht , Danish Pruthi

Powerful predictive AI systems have demonstrated great potential in augmenting human decision making. Recent empirical work has argued that the vision for optimal human-AI collaboration requires 'appropriate reliance' of humans on AI…

Artificial Intelligence · Computer Science 2024-09-24 Gaole He , Abri Bharos , Ujwal Gadiraju

Large language models trained on human feedback may suppress fraud warnings when investors arrive already persuaded of a fraudulent opportunity. We tested this in a preregistered experiment across seven leading LLMs and twelve investment…

Artificial Intelligence · Computer Science 2026-04-24 Nattavudh Powdthavee

Humans are black boxes -- we cannot observe their neural processes, yet society functions by evaluating verifiable arguments. AI explainability should follow this principle: stakeholders need verifiable reasoning chains, not mechanistic…

Machine Learning · Computer Science 2025-10-07 Ege Cakar , Per Ola Kristensson

As the capabilities of large machine learning models continue to grow, and as the autonomy afforded to such models continues to expand, the spectre of a new adversary looms: the models themselves. The threat that a model might behave in a…

Machine Learning · Computer Science 2023-07-27 Andres Carranza , Dhruv Pai , Rylan Schaeffer , Arnuv Tandon , Sanmi Koyejo

To support human decision making with machine learning models, we often need to elucidate patterns embedded in the models that are unsalient, unknown, or counterintuitive to humans. While existing approaches focus on explaining machine…

Human-Computer Interaction · Computer Science 2020-01-17 Vivian Lai , Han Liu , Chenhao Tan

We investigate problems in penalized $M$-estimation, inspired by applications in machine learning debugging. Data are collected from two pools, one containing data with possibly contaminated labels, and the other which is known to contain…

Machine Learning · Computer Science 2021-08-11 Xiaomin Zhang , Xiaojin Zhu , Po-Ling Loh

With the recent advent of Large Language Models (LLMs), such as ChatGPT from OpenAI, BARD from Google, Llama2 from Meta, and Claude from Anthropic AI, gain widespread use, ensuring their security and robustness is critical. The widespread…

Human-Computer Interaction · Computer Science 2023-11-28 Sonali Singh , Faranak Abri , Akbar Siami Namin

We show that language models' activations linearly encode when information was learned during training. Our setup involves creating a model with a known training order by sequentially fine-tuning Llama-3.2-1B on six disjoint but otherwise…

Machine Learning · Computer Science 2025-09-23 Dmitrii Krasheninnikov , Richard E. Turner , David Krueger

This study investigates the ability of multimodal Large Language Models (LLMs) to identify and interpret misleading visualizations, and recognize these observations along with their underlying causes and potential intentionality. Our…

Human-Computer Interaction · Computer Science 2026-04-02 Graziano Blasilli , Marco Angelini

Academic integrity continues to face the persistent challenge of examination cheating. Traditional invigilation relies on human observation, which is inefficient, costly, and prone to errors at scale. Although some existing AI-powered…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Van-Truong Le , Le-Khanh Nguyen , Trong-Doanh Nguyen

A common assumption in the social learning literature is that agents exchange information in an unselfish manner. In this work, we consider the scenario where a subset of agents aims at deceiving the network, meaning they aim at driving the…

Systems and Control · Electrical Eng. & Systems 2021-03-30 Konstantinos Ntemos , Virginia Bordignon , Stefan Vlaski , Ali H. Sayed

Developing effective world models is crucial for creating artificial agents that can reason about and navigate complex environments. In this paper, we investigate a deep supervision technique for encouraging the development of a world model…

Artificial Intelligence · Computer Science 2025-04-08 Andrii Zahorodnii

Large Language Models (LLMs) perform impressively well in various applications. However, the potential for misuse of these models in activities such as plagiarism, generating fake news, and spamming has raised concern about their…

Computation and Language · Computer Science 2025-01-20 Vinu Sankar Sadasivan , Aounon Kumar , Sriram Balasubramanian , Wenxiao Wang , Soheil Feizi

To reliably assist human decision-making, LLMs must maintain factual internal beliefs against misleading injections. While current models resist explicit misinformation, we uncover a fundamental vulnerability to sophisticated,…

Computation and Language · Computer Science 2026-01-12 Herun Wan , Jiaying Wu , Minnan Luo , Fanxiao Li , Zhi Zeng , Min-Yen Kan

AI agents that execute tasks via tool calls frequently hallucinate results - fabricating tool executions, misstating output counts, or presenting inferences as facts. Recent approaches to verifiable AI inference rely on zero-knowledge…

Cryptography and Security · Computer Science 2026-03-12 Abhinaba Basu

Hint-based faithfulness evaluations have established that Large Reasoning Models (LRMs) may not say what they think: they do not always volunteer information about how key parts of the input (e.g. answer hints) influence their reasoning.…

Artificial Intelligence · Computer Science 2026-04-22 William Walden , Miriam Wanner

Large Language Model (LLM)-based agents are increasingly used as autonomous subordinates that carry out tasks for users. This raises the question of whether they may also engage in deception, similar to how individuals in human…

‹ Prev 1 8 9 10 Next ›