English

Delving into adversarial attacks on deep policies

Machine Learning 2017-05-19 v1 Machine Learning

Abstract

Adversarial examples have been shown to exist for a variety of deep learning architectures. Deep reinforcement learning has shown promising results on training agent policies directly on raw inputs such as image pixels. In this paper we present a novel study into adversarial attacks on deep reinforcement learning polices. We compare the effectiveness of the attacks using adversarial examples vs. random noise. We present a novel method for reducing the number of times adversarial examples need to be injected for a successful attack, based on the value function. We further explore how re-training on random noise and FGSM perturbations affects the resilience against adversarial examples.

Keywords

Cite

@article{arxiv.1705.06452,
  title  = {Delving into adversarial attacks on deep policies},
  author = {Jernej Kos and Dawn Song},
  journal= {arXiv preprint arXiv:1705.06452},
  year   = {2017}
}

Comments

ICLR 2017 Workshop

R2 v1 2026-06-22T19:50:48.035Z