English

Hidden in Plain Sight -- Class Competition Focuses Attribution Maps

Computer Vision and Pattern Recognition 2026-02-06 v3 Machine Learning

Abstract

Attribution methods reveal which input features a neural network uses for a prediction, adding transparency to their decisions. A common problem is that these attributions seem unspecific, highlighting both important and irrelevant features. We revisit the common attribution pipeline and observe that using logits as attribution target is a main cause of this phenomenon. We show that the solution is in plain sight: considering distributions of attributions over multiple classes using existing attribution methods yields specific and fine-grained attributions. On common benchmarks, including the grid-pointing game and randomization-based sanity checks, this improves the ability of 18 attribution methods across 7 architectures up to 2x, agnostic to model architecture.

Keywords

Cite

@article{arxiv.2503.07346,
  title  = {Hidden in Plain Sight -- Class Competition Focuses Attribution Maps},
  author = {Nils Philipp Walter and Jilles Vreeken and Jonas Fischer},
  journal= {arXiv preprint arXiv:2503.07346},
  year   = {2026}
}
R2 v1 2026-06-28T22:14:05.472Z