解释是有用的:通过惩罚解释使神经网络与先验知识对齐
机器学习
2020-10-09 v4 计算机视觉与模式识别
机器学习
摘要
要使深度学习模型的解释有效,它必须既提供对模型的洞察,又提出相应的行动以实现某个目标。已有大量提出的可解释深度学习方法常止步于第一步,为从业者提供对模型的洞察,却无据此行动之途径。本文提出上下文分解解释惩罚(contextual decomposition explanation penalization, CDEP)方法,使从业者能够利用现有解释方法来提升深度学习模型的预测精度。具体而言,当得知模型错误地赋予某些特征重要性时,CDEP 使从业者能够通过直接正则化所提供的解释来纠正这些错误。利用上下文分解(contextual decomposition, CD)(Murdoch 等,2018)提供的解释,我们展示了该方法在一系列合成与真实数据集上提升性能的能力。
引用
@article{arxiv.1909.13584,
title = {Interpretations are useful: penalizing explanations to align neural networks with prior knowledge},
author = {Laura Rieger and Chandan Singh and W. James Murdoch and Bin Yu},
journal= {arXiv preprint arXiv:1909.13584},
year = {2020}
}
备注
18 pages; published in ICML2020; Erratum: numbers in table 1 were too high (now corrected) with the trend remaining the same