English

Poisoning Deep Reinforcement Learning Agents with In-Distribution Triggers

Machine Learning 2021-06-16 v1 Cryptography and Security

Abstract

In this paper, we propose a new data poisoning attack and apply it to deep reinforcement learning agents. Our attack centers on what we call in-distribution triggers, which are triggers native to the data distributions the model will be trained on and deployed in. We outline a simple procedure for embedding these, and other, triggers in deep reinforcement learning agents following a multi-task learning paradigm, and demonstrate in three common reinforcement learning environments. We believe that this work has important implications for the security of deep learning models.

Keywords

Cite

@article{arxiv.2106.07798,
  title  = {Poisoning Deep Reinforcement Learning Agents with In-Distribution Triggers},
  author = {Chace Ashcraft and Kiran Karra},
  journal= {arXiv preprint arXiv:2106.07798},
  year   = {2021}
}

Comments

4 pages, 1 figure, Published at ICLR 2021 Workshop on Security and Safety in Machine Learning Systems

R2 v1 2026-06-24T03:12:03.218Z