中文

模型可解释性中人类评估的承诺与危险

人工智能 2019-10-31 v2 机器学习 机器学习

摘要

透明性、用户信任与人类理解是可解释机器学习的常见伦理动机。为支持这些目标,研究者利用人类和真实世界应用来评估模型解释性能。这本身在人工智能许多领域就构成一项挑战。在本立场论文中,我们提出描述性解释与说服性解释之间的区分。我们讨论一种推理,其表明功能性可解释性可能与认知功能和用户偏好相关。若事实的确如此,使用功能性度量进行评估与优化可能会延续解释中威胁透明性的隐性认知偏见。最后,我们提出两个潜在研究方向以消歧认知功能与解释模型,保留对准确性与可解释性之间权衡的控制。

关键词

引用

@article{arxiv.1711.07414,
  title  = {The Promise and Peril of Human Evaluation for Model Interpretability},
  author = {Bernease Herman},
  journal= {arXiv preprint arXiv:1711.07414},
  year   = {2019}
}

备注

Presented at NIPS 2017 Symposium on Interpretable Machine Learning. I'm not happy with the writing and presentation of these ideas and hope to submit an updated and extended version in 2020