English

The Intriguing Properties of Model Explanations

Machine Learning 2018-01-31 v1 Artificial Intelligence

Abstract

Linear approximations to the decision boundary of a complex model have become one of the most popular tools for interpreting predictions. In this paper, we study such linear explanations produced either post-hoc by a few recent methods or generated along with predictions with contextual explanation networks (CENs). We focus on two questions: (i) whether linear explanations are always consistent or can be misleading, and (ii) when integrated into the prediction process, whether and how explanations affect the performance of the model. Our analysis sheds more light on certain properties of explanations produced by different methods and suggests that learning models that explain and predict jointly is often advantageous.

Keywords

Cite

@article{arxiv.1801.09808,
  title  = {The Intriguing Properties of Model Explanations},
  author = {Maruan Al-Shedivat and Avinava Dubey and Eric P. Xing},
  journal= {arXiv preprint arXiv:1801.09808},
  year   = {2018}
}

Comments

Interpretable ML Symposium, NIPS 2017

R2 v1 2026-06-23T00:02:39.444Z