中文

CLIF:面向透明瓶颈模型的概念级影响函数

计算与语言 2026-05-26 v2

摘要

近年来,深度学习模型的黑箱特性限制了其在医疗诊断和金融等高风险领域的应用,在这些领域中,可解释性至关重要。为解决此问题,我们提出了一种新颖的方法,利用影响函数在样本和概念两个层面增强自然语言处理模型的可解释性。在 CEBaB 和 Yelp 数据集上的实验表明,影响函数能有效识别对模型预测影响最大的训练样本,无论是有益还是有害的。通过调整这些样本的标签和权重,我们证明模型性能可以在不重新训练的情况下恢复到基线水平,证实了影响函数在高效数据调试方面的价值。此外,我们的概念级分析识别出概念瓶颈模型(Concept Bottleneck Models, CBM)中显著影响预测的关键概念。修改这些概念可观察到模型行为的改变,从而为决策过程提供清晰的洞察。

关键词

引用

@article{arxiv.2605.19848,
  title  = {CLIF: Concept-Level Influence Functions for Transparent Bottleneck Models},
  author = {Yike Sun and Mingkun Xu and Mu You and Zhongzhi He and Henghua Shen and Zehan Tan and Derek F. Wong and Tao Fang},
  journal= {arXiv preprint arXiv:2605.19848},
  year   = {2026}
}

备注

A critical theoretical error invalidates the main results. The independence assumption on concept representations and gradients (Section 3.2, Eq.7) is incorrect, breaking the influence estimation in nonlinear bottleneck layers. This flaw undermines all empirical claims in Sections 4-5. The authors withdraw to prevent dissemination of incorrect findings