English

Breaking certified defenses: Semantic adversarial examples with spoofed robustness certificates

Machine Learning 2020-03-20 v1 Machine Learning

Abstract

To deflect adversarial attacks, a range of "certified" classifiers have been proposed. In addition to labeling an image, certified classifiers produce (when possible) a certificate guaranteeing that the input image is not an p\ell_p-bounded adversarial example. We present a new attack that exploits not only the labelling function of a classifier, but also the certificate generator. The proposed method applies large perturbations that place images far from a class boundary while maintaining the imperceptibility property of adversarial examples. The proposed "Shadow Attack" causes certifiably robust networks to mislabel an image and simultaneously produce a "spoofed" certificate of robustness.

Keywords

Cite

@article{arxiv.2003.08937,
  title  = {Breaking certified defenses: Semantic adversarial examples with spoofed robustness certificates},
  author = {Amin Ghiasi and Ali Shafahi and Tom Goldstein},
  journal= {arXiv preprint arXiv:2003.08937},
  year   = {2020}
}
R2 v1 2026-06-23T14:20:34.684Z