中文

解释是有用的:通过惩罚解释使神经网络与先验知识对齐

机器学习 2020-10-09 v4 计算机视觉与模式识别 机器学习

摘要

要使深度学习模型的解释有效,它必须既提供对模型的洞察,又提出相应的行动以实现某个目标。已有大量提出的可解释深度学习方法常止步于第一步,为从业者提供对模型的洞察,却无据此行动之途径。本文提出上下文分解解释惩罚(contextual decomposition explanation penalization, CDEP)方法,使从业者能够利用现有解释方法来提升深度学习模型的预测精度。具体而言,当得知模型错误地赋予某些特征重要性时,CDEP 使从业者能够通过直接正则化所提供的解释来纠正这些错误。利用上下文分解(contextual decomposition, CD)(Murdoch 等,2018)提供的解释,我们展示了该方法在一系列合成与真实数据集上提升性能的能力。

关键词

引用

@article{arxiv.1909.13584,
  title  = {Interpretations are useful: penalizing explanations to align neural networks with prior knowledge},
  author = {Laura Rieger and Chandan Singh and W. James Murdoch and Bin Yu},
  journal= {arXiv preprint arXiv:1909.13584},
  year   = {2020}
}

备注

18 pages; published in ICML2020; Erratum: numbers in table 1 were too high (now corrected) with the trend remaining the same