English

An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L

Machine Learning 2024-12-17 v4 Artificial Intelligence

Abstract

Prior work suggests that language models manage the limited bandwidth of the residual stream through a "memory management" mechanism, where certain attention heads and MLP layers clear residual stream directions set by earlier layers. Our study provides concrete evidence for this erasure phenomenon in a 4-layer transformer, identifying heads that consistently remove the output of earlier heads. We further demonstrate that direct logit attribution (DLA), a common technique for interpreting the output of intermediate transformer layers, can show misleading results by not accounting for erasure.

Keywords

Cite

@article{arxiv.2310.07325,
  title  = {An Adversarial Example for Direct Logit Attribution: Memory Management in GELU-4L},
  author = {Jett Janiak and Can Rager and James Dao and Yeu-Tong Lau},
  journal= {arXiv preprint arXiv:2310.07325},
  year   = {2024}
}
R2 v1 2026-06-28T12:47:07.474Z