English

Measuring Explainer Stability via Attribution Separability

Machine Learning 2026-08-03 v1 Artificial Intelligence

Abstract

Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can produce variable attribution scores due to stochastic components in their definition. In this paper, we propose a distribution-based framework to capture the stability of attribution scores. In particular, our approach allows to understand the degree of separability in the ranked attribution vector and obtain the largest index for which a feature ranking remains reliable. We further extend this framework to compare AMs based on the robustness of their rankings across a dataset. Through experiments, we demonstrate how to apply our method to evaluate explainer stability. Overall, our approach provides a complementary criterion for evaluating the stability of AMs.

Cite

@article{arxiv.2608.02697,
  title  = {Measuring Explainer Stability via Attribution Separability},
  author = {Eddie Conti and Álvaro Parafita and Axel Brando},
  journal= {arXiv preprint arXiv:2608.02697},
  year   = {2026}
}

Comments

Accepted at EXPLAINS 2026 Conference