English

Actionable Interpretability Must Be Defined in Terms of Symmetries

Artificial Intelligence 2026-01-30 v3 Machine Learning Neural and Evolutionary Computing

Abstract

This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpretability can be formally tested or designed for. We posit that actionable definitions of interpretability must be formulated in terms of *symmetries* that inform model design and lead to testable conditions. Under a probabilistic view, we hypothesise that four symmetries (inference equivariance, information invariance, concept-closure invariance, and structural invariance) suffice to (i) formalise interpretable models as a subclass of probabilistic models, (ii) yield a unified formulation of interpretable inference (e.g., alignment, interventions, and counterfactuals) as a form of Bayesian inversion, and (iii) provide a formal framework to verify compliance with safety standards and regulations.

Keywords

Cite

@article{arxiv.2601.12913,
  title  = {Actionable Interpretability Must Be Defined in Terms of Symmetries},
  author = {Pietro Barbiero and Mateo Espinosa Zarlenga and Francesco Giannini and Alberto Termine and Filippo Bonchi and Mateja Jamnik and Giuseppe Marra},
  journal= {arXiv preprint arXiv:2601.12913},
  year   = {2026}
}