知识显微镜:特征作为比神经元更好的分析镜头
计算与语言
2025-02-28 v2
摘要
先前的研究主要利用MLP神经元作为理解语言模型(LM)中事实知识机制的分析单元;然而,神经元存在多义性,导致知识表达有限且解释性差。在本文中,我们首先进行初步实验以验证稀疏自动编码器(SAE)能够有效地将神经元分解为特征,这些特征作为替代的分析单元。 established after that, our core findings reveal three key advantages of features over neurons: (1) Features exhibit stronger influence on knowledge expression and superior interpretability. (2) Features demonstrate enhanced monosemanticity, showing distinct activation patterns between related and unrelated facts. (3) Features achieve better privacy protection than neurons, demonstrated through our proposed FeatureEdit method, which significantly outperforms existing neuron-based approaches in erasing privacy-sensitive information from LMs.Code and dataset will be available.
引用
@article{arxiv.2502.12483,
title = {The Knowledge Microscope: Features as Better Analytical Lenses than Neurons},
author = {Yuheng Chen and Pengfei Cao and Kang Liu and Jun Zhao},
journal= {arXiv preprint arXiv:2502.12483},
year = {2025}
}
备注
ARR February UnderReview