基于谱分析的解释质量结构探索
摘要
随着 machine learning models being increasingly considered for high-stakes domains, effective explanation methods 至关重要 to ensure that their prediction strategies are transparent to the user。多年来,已提出 numerous metrics to assess quality of explanations。然而,其 practical applicability 仍不清晰,尤其是 due to a limited understanding of which specific aspects each metric rewards。本文提出 a new framework based on spectral analysis of explanation outcomes,系统地 capture the multifaceted properties of different explanation techniques。我们的分析揭示了两种 distinct factors of explanation quality—stability and target sensitivity—通过 spectral decomposition 直接可见。MNIST 和 ImageNet 上的实验表明,流行的 evaluation techniques(如 pixel-flipping, entropy)部分 capture 这些 factors 之间的 trade-offs。总体而言,我们的 framework 提供了 understanding explanation quality 的基础,指导 develop more reliable techniques for evaluating explanations。
引用
@article{arxiv.2504.08553,
title = {Uncovering the Structure of Explanation Quality with Spectral Analysis},
author = {Johannes Maeß and Grégoire Montavon and Shinichi Nakajima and Klaus-Robert Müller and Thomas Schnake},
journal= {arXiv preprint arXiv:2504.08553},
year = {2025}
}
备注
14 pages, 5 figures, Accepted at XAI World Conference 2025