English

Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partition

Artificial Intelligence 2025-11-12 v1

Abstract

Attribution-based explanation techniques capture key patterns to enhance visual interpretability; however, these patterns often lack the granularity needed for insight in fine-grained tasks, particularly in cases of model misclassification, where explanations may be insufficiently detailed. To address this limitation, we propose a fine-grained counterfactual explanation framework that generates both object-level and part-level interpretability, addressing two fundamental questions: (1) which fine-grained features contribute to model misclassification, and (2) where dominant local features influence counterfactual adjustments. Our approach yields explainable counterfactuals in a non-generative manner by quantifying similarity and weighting component contributions within regions of interest between correctly classified and misclassified samples. Furthermore, we introduce a saliency partition module grounded in Shapley value contributions, isolating features with region-specific relevance. Extensive experiments demonstrate the superiority of our approach in capturing more granular, intuitively meaningful regions, surpassing fine-grained methods.

Keywords

Cite

@article{arxiv.2511.07974,
  title  = {Towards Fine-Grained Interpretability: Counterfactual Explanations for Misclassification with Saliency Partition},
  author = {Lintong Zhang and Kang Yin and Seong-Whan Lee},
  journal= {arXiv preprint arXiv:2511.07974},
  year   = {2025}
}
R2 v1 2026-07-01T07:31:32.097Z