English

Are Perceptually-Aligned Gradients a General Property of Robust Classifiers?

Machine Learning 2019-10-24 v2 Computer Vision and Pattern Recognition Machine Learning

Abstract

For a standard convolutional neural network, optimizing over the input pixels to maximize the score of some target class will generally produce a grainy-looking version of the original image. However, Santurkar et al. (2019) demonstrated that for adversarially-trained neural networks, this optimization produces images that uncannily resemble the target class. In this paper, we show that these "perceptually-aligned gradients" also occur under randomized smoothing, an alternative means of constructing adversarially-robust classifiers. Our finding supports the hypothesis that perceptually-aligned gradients may be a general property of robust classifiers. We hope that our results will inspire research aimed at explaining this link between perceptually-aligned gradients and adversarial robustness.

Keywords

Cite

@article{arxiv.1910.08640,
  title  = {Are Perceptually-Aligned Gradients a General Property of Robust Classifiers?},
  author = {Simran Kaur and Jeremy Cohen and Zachary C. Lipton},
  journal= {arXiv preprint arXiv:1910.08640},
  year   = {2019}
}

Comments

To appear in the "Science Meets Engineering of Deep Learning" Workshop at NeurIPS 2019

R2 v1 2026-06-23T11:48:17.259Z