English

How Faithful are Self-Explainable GNNs?

Machine Learning 2023-08-30 v1

Abstract

Self-explainable deep neural networks are a recent class of models that can output ante-hoc local explanations that are faithful to the model's reasoning, and as such represent a step forward toward filling the gap between expressiveness and interpretability. Self-explainable graph neural networks (GNNs) aim at achieving the same in the context of graph data. This begs the question: do these models fulfill their implicit guarantees in terms of faithfulness? In this extended abstract, we analyze the faithfulness of several self-explainable GNNs using different measures of faithfulness, identify several limitations -- both in the models themselves and in the evaluation metrics -- and outline possible ways forward.

Keywords

Cite

@article{arxiv.2308.15096,
  title  = {How Faithful are Self-Explainable GNNs?},
  author = {Marc Christiansen and Lea Villadsen and Zhiqiang Zhong and Stefano Teso and Davide Mottin},
  journal= {arXiv preprint arXiv:2308.15096},
  year   = {2023}
}
R2 v1 2026-06-28T12:07:02.101Z