English

Adversarial Attacks on the Interpretation of Neuron Activation Maximization

Machine Learning 2023-06-14 v1 Computer Vision and Pattern Recognition

Abstract

The internal functional behavior of trained Deep Neural Networks is notoriously difficult to interpret. Activation-maximization approaches are one set of techniques used to interpret and analyze trained deep-learning models. These consist in finding inputs that maximally activate a given neuron or feature map. These inputs can be selected from a data set or obtained by optimization. However, interpretability methods may be subject to being deceived. In this work, we consider the concept of an adversary manipulating a model for the purpose of deceiving the interpretation. We propose an optimization framework for performing this manipulation and demonstrate a number of ways that popular activation-maximization interpretation techniques associated with CNNs can be manipulated to change the interpretations, shedding light on the reliability of these methods.

Keywords

Cite

@article{arxiv.2306.07397,
  title  = {Adversarial Attacks on the Interpretation of Neuron Activation Maximization},
  author = {Geraldin Nanfack and Alexander Fulleringer and Jonathan Marty and Michael Eickenberg and Eugene Belilovsky},
  journal= {arXiv preprint arXiv:2306.07397},
  year   = {2023}
}
R2 v1 2026-06-28T11:03:22.863Z