特征归因稳定性套件: 后向归因方法有多稳定?
计算机视觉与模式识别
2026-04-06 v1 人工智能
机器学习
摘要
后向特征归因方法在安全关键的视觉系统中广泛部署, 但其在现实输入扰动下的稳定性仍未得到良好 characterize。现有指标在加性噪声下评估解释, 将稳定性简化为单个标量, 并未以预测保持为条件, 导致解释脆弱性与模型灵敏性混为一体。我们引入特征归因稳定性套件(FASS), 一个强制执行预测不变性过滤、分解为三个互补指标: 结构相似性、秩相关性和 top-k Jaccard 重叠, 并在几何、光度和压缩扰动下进行评估。我们在四种归因方法(Integrated Gradients、GradientSHAP、Grad-CAM、LIME)上, 跨越四种架构和三个数据集-ImageNet-1K、MS COCO 和 CIFAR-10, 评估 FASS, 发现稳定性估计在扰动系列和预测不变性过滤方面具有关键依赖性。几何扰动暴露出显著大于光度变化的归因不稳定性, 在未对预测保持进行条件控制时, 最多有 99% 的评估对涉及改变预测。在受控评估下, 我们观察到一致的方法水平趋势, Grad-CAM 在数据集之间实现最高的稳定性。
引用
@article{arxiv.2604.02532,
title = {Feature Attribution Stability Suite: How Stable Are Post-Hoc Attributions?},
author = {Kamalasankari Subramaniakuppusamy and Jugal Gajjar},
journal= {arXiv preprint arXiv:2604.02532},
year = {2026}
}
备注
Accepted in the proceedings track of XAI4CV Workshop at CVPR 2026. It has 2 images, 5 tables, 6 equations, and 35 references in the main paper and 12 figures, 15 tables, and 3 references in the supplementary material