English
Related papers

Related papers: Experiments with Detecting and Mitigating AI Decep…

200 papers

Audio has become an increasingly crucial biometric modality due to its ability to provide an intuitive way for humans to interact with machines. It is currently being used for a range of applications, including person authentication to…

Sound · Computer Science 2023-07-14 Rishabh Ranjan , Mayank Vatsa , Richa Singh

In this paper, we study the use of deception for strategic planning in adversarial environments. We model the interaction between the agent (player 1) and the adversary (player 2) as a two-player concurrent game in which the adversary has…

Computer Science and Game Theory · Computer Science 2020-08-03 Lening Li , Haoxiang Ma , Abhishek N. Kulkarni , Jie Fu

Lie detection is considered a concern for everyone in their day to day life given its impact on human interactions. Thus, people normally pay attention to both what their interlocutors are saying and also to their visual appearances,…

Computer Vision and Pattern Recognition · Computer Science 2021-07-01 Nuria Rodriguez-Diaz , Decky Aspandi , Federico Sukno , Xavier Binefa

In this work, we study a class of deception planning problems in which an agent aims to alter a security monitoring system's sensor readings so as to disguise its adversarial itinerary as an allowed itinerary in the environment. The…

Robotics · Computer Science 2024-12-04 Hazhar Rahmani , Arash Ahadi , Jie Fu

We demonstrate how AI agents can coordinate to deceive oversight systems using automated interpretability of neural networks. Using sparse autoencoders (SAEs) as our experimental framework, we show that language models (Llama, DeepSeek R1,…

Artificial Intelligence · Computer Science 2025-04-11 Simon Lermen , Mateusz Dziemian , Natalia Pérez-Campanero Antolín

The conversation around artificial intelligence (AI) often focuses on safety, transparency, accountability, alignment, and responsibility. However, AI security (i.e., the safeguarding of data, models, and pipelines from adversarial…

Cryptography and Security · Computer Science 2025-04-24 Krti Tallam

The implementation of agentic AI systems has the potential of providing more helpful AI systems in a variety of applications. These systems work autonomously towards a defined goal with reduced external control. Despite their potential, one…

Artificial Intelligence · Computer Science 2025-11-13 Niclas Flehmig , Mary Ann Lundteigen , Shen Yin

We propose an information-theoretic formalization of the distinction between two fundamental AI safety failure modes: deceptive alignment and goal drift. While both can lead to systems that appear misaligned, we demonstrate that they…

Artificial Intelligence · Computer Science 2026-03-31 Robin Young

The act of bluffing confounds game designers to this day. The very nature of bluffing is even open for debate, adding further complication to the process of creating intelligent virtual players that can bluff, and hence play, realistically.…

Artificial Intelligence · Computer Science 2007-05-23 Evan Hurwitz , Tshilidzi Marwala

Multi-agent systems leverage advanced AI models as autonomous agents that interact, cooperate, or compete to complete complex tasks across applications such as robotics and traffic management. Despite their growing importance, safety in…

Multiagent Systems · Computer Science 2025-05-28 Falong Fan , Xi Li

In this paper, we study the problem of deceptive reinforcement learning to preserve the privacy of a reward function. Reinforcement learning is the problem of finding a behaviour policy based on rewards received from exploratory behaviour.…

Machine Learning · Computer Science 2021-02-08 Zhengshang Liu , Yue Yang , Tim Miller , Peta Masters

AI agents are increasingly deployed and used to make automated decisions that affect our lives on a daily basis. It is imperative to ensure that these systems embed ethical principles and respect human values. We focus on how we can attest…

Artificial Intelligence · Computer Science 2019-09-11 Xavier Ferrer Aran , Jose M. Such , Natalia Criado

Deceptive patterns are design practices embedded in digital platforms to manipulate users, representing a widespread and long-standing issue in the web and mobile software development industry. Legislative actions highlight the urgency of…

Cryptography and Security · Computer Science 2024-02-07 Zewei Shi , Ruoxi Sun , Jieshan Chen , Jiamou Sun , Minhui Xue

Distributed Denial of Service attacks represent an active cybersecurity research problem. Recent research shifted from static rule-based defenses towards AI-based detection and mitigation. This comprehensive survey covers several key…

Cryptography and Security · Computer Science 2026-03-20 Alexandru Apostu , Silviu Gheorghe , Andrei Hîji , Nicolae Cleju , Andrei Pătraşcu , Cristian Rusu , Radu Ionescu , Paul Irofti

Agents operating in physical environments need to be able to handle delays in the input and output signals since neither data transmission nor sensing or actuating the environment are instantaneous. Shields are correct-by-construction…

Artificial Intelligence · Computer Science 2023-07-06 Filip Cano Córdoba , Alexander Palmisano , Martin Fränzle , Roderick Bloem , Bettina Könighofer

Background: Deception detection is a prevalent problem for security practitioners. With a need for more large-scale approaches, automated methods using machine learning have gained traction. However, detection performance still implies…

Computation and Language · Computer Science 2020-03-31 Bennett Kleinberg , Bruno Verschuere

Large language models (LLMs) aligned for safety through techniques like reinforcement learning from human feedback (RLHF) often exhibit emergent deceptive behaviors, where outputs appear compliant but subtly mislead or omit critical…

Machine Learning · Computer Science 2025-07-15 Santhosh Kumar Ravindran

Mitigating bias in algorithmic systems is a critical issue drawing attention across communities within the information and computer sciences. Given the complexity of the problem and the involvement of multiple stakeholders -- including…

Autonomous agents trained via reinforcement learning present numerous safety concerns: reward hacking, negative side effects, and unsafe exploration, among others. In the context of near-future autonomous agents, operating in environments…

Artificial Intelligence · Computer Science 2019-02-20 Christopher Frye , Ilya Feige

Artificial Intelligence (AI) has been used extensively in automatic decision making in a broad variety of scenarios, ranging from credit ratings for loans to recommendations of movies. Traditional design guidelines for AI models focus…

Artificial Intelligence · Computer Science 2018-09-27 Marisa Vasconcelos , Carlos Cardonha , Bernardo Gonçalves
‹ Prev 1 4 5 6 7 8 10 Next ›