characterize LLM残差流中的稳定区域
机器学习
2024-11-19 v4
摘要
我们识别了变换器模型残差流中的稳定区域,其中模型的输出对小的激活变化不敏感,但在区域边界处表现出高度敏感性。这些区域在训练期间出现,并且随着训练进程的推进或模型规模的增加而变得更加明确。这些区域似乎远大于此前研究的多面体。我们的分析表明,这些稳定区域与语义区分相吻合,即相似的提示在同一区域内聚类,来自同一区域的激活会导致相似的下一个标记预测。本工作为理解神经网络的复杂性、阐明训练动力学以及推进可解释性提供了有前景的研究方向。
引用
@article{arxiv.2409.17113,
title = {Characterizing stable regions in the residual stream of LLMs},
author = {Jett Janiak and Jacek Karwowski and Chatrik Singh Mangat and Giorgi Giglemiani and Nora Petrova and Stefan Heimersheim},
journal= {arXiv preprint arXiv:2409.17113},
year = {2024}
}
备注
Presented at the Scientific Methods for Understanding Deep Learning (SciForDL) workshop at NeurIPS 2024