中文

RAVEL:评估语言模型表示解缠的可解释性方法

计算与语言 2024-08-28 v2 机器学习

摘要

单个神经元参与多个高层概念的表示。不同的可解释性方法在多大程度上能成功解缠这些角色?为了帮助回答这个问题,我们引入了 RAVEL(解决语言模型中的属性 - 值纠缠),这是一个数据集,能够对各种现有的可解释性方法进行紧密控制的定量比较。我们利用由此产生的概念框架定义了新的多任务分布式对齐搜索 (MDAS) 方法,该方法使我们能够找到满足多个因果标准的分布式表示。以 Llama2-7B 为目标语言模型,MDAS 在 RAVEL 上取得了最先进的结果,证明了超越神经元级分析以识别跨激活分布的特征的重要性。我们在 https://github.com/explanare/ravel 发布了我们的基准测试。

关键词

引用

@article{arxiv.2402.17700,
  title  = {RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations},
  author = {Jing Huang and Zhengxuan Wu and Christopher Potts and Mor Geva and Atticus Geiger},
  journal= {arXiv preprint arXiv:2402.17700},
  year   = {2024}
}

备注

Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024)