English

The Role of Syntactic Span Preferences in Post-Hoc Explanation Disagreement

Computation and Language 2024-03-29 v1 Artificial Intelligence

Abstract

Post-hoc explanation methods are an important tool for increasing model transparency for users. Unfortunately, the currently used methods for attributing token importance often yield diverging patterns. In this work, we study potential sources of disagreement across methods from a linguistic perspective. We find that different methods systematically select different classes of words and that methods that agree most with other methods and with humans display similar linguistic preferences. Token-level differences between methods are smoothed out if we compare them on the syntactic span level. We also find higher agreement across methods by estimating the most important spans dynamically instead of relying on a fixed subset of size kk. We systematically investigate the interaction between kk and spans and propose an improved configuration for selecting important tokens.

Keywords

Cite

@article{arxiv.2403.19424,
  title  = {The Role of Syntactic Span Preferences in Post-Hoc Explanation Disagreement},
  author = {Jonathan Kamp and Lisa Beinborn and Antske Fokkens},
  journal= {arXiv preprint arXiv:2403.19424},
  year   = {2024}
}

Comments

Long paper accepted to LREC-Coling 2024 main conference. Please cite the conference proceedings version when available