Computation and Language · Computer Science
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
Stephen Zhang, Mustafa Khan, Vardan Papyan
2025-09-23
Machine Learning · Computer Science
Layer by Layer: Uncovering Hidden Representations in Language Models
Oscar Skean, Md Rifat Arefin, Dan Zhao, Niket Patel +3
2025-06-17
Machine Learning · Computer Science
Attention Sinks and Compression Valleys in LLMs are Two Sides of the Same Coin
Enrique Queipo-de-Llano, Álvaro Arroyo, Federico Barbero, Xiaowen Dong +3
2026-02-11
Computation and Language · Computer Science
Attention Lens: A Tool for Mechanistically Interpreting the Attention Head Information Retrieval Mechanism
Mansi Sakarvadia, Arham Khan, Aswathy Ajith, Daniel Grzenda +4
2023-10-26
Computer Vision and Pattern Recognition · Computer Science
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
Jiho Choi, Jaemin Kim, Sanghwan Kim, Seunghoon Hong +1
2026-04-07
Computation and Language · Computer Science
When Attention Sink Emerges in Language Models: An Empirical View
Xiangming Gu, Tianyu Pang, Chao Du, Qian Liu +4
2025-03-04
Machine Learning · Computer Science
Interpreting the Repeated Token Phenomenon in Large Language Models
Itay Yona, Ilia Shumailov, Jamie Hayes, Federico Barbero +1
2025-03-13
Computation and Language · Computer Science
Why do LLMs attend to the first token?
Federico Barbero, Álvaro Arroyo, Xiangming Gu, Christos Perivolaropoulos +3
2025-08-06
Computation and Language · Computer Science
Attention Sinks in Diffusion Language Models
Maximo Eduardo Rulli, Simone Petruzzi, Edoardo Michielon, Fabrizio Silvestri +2
2025-12-11
Computer Vision and Pattern Recognition · Computer Science
Tinted Frames: Question Framing Blinds Vision-Language Models
Wan-Cyuan Fan, Jiayun Luo, Declan Kutscher, Leonid Sigal +1
2026-03-23
Computation and Language · Computer Science
Residual Connections and the Causal Shift: Uncovering a Structural Misalignment in Transformers
Jonathan Lys, Vincent Gripon, Bastien Pasdeloup, Axel Marmoret +3
2026-03-03
Machine Learning · Computer Science
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
Oscar Skean, Md Rifat Arefin, Yann LeCun, Ravid Shwartz-Ziv
2024-12-13
Machine Learning · Computer Science
Spectral Insights into Data-Oblivious Critical Layers in Large Language Models
Xuyuan Liu, Lei Hsiung, Yaoqing Yang, Yujun Yan
2025-06-06
Computer Vision and Pattern Recognition · Computer Science
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
Jiayun Luo, Wan-Cyuan Fan, Lyuyang Wang, Xiangteng He +3
2025-10-10
Computation and Language · Computer Science
LLMs Explain't: A Post-Mortem on Semantic Interpretability in Transformer Models
Alhassan Abdelhalim, Janick Edinger, Sören Laue, Michaela Regneri
2026-02-02
Computer Vision and Pattern Recognition · Computer Science
Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens
Zhangqi Jiang, Junkai Chen, Beier Zhu, Tingjin Luo +2
2025-04-02
Computer Vision and Pattern Recognition · Computer Science
See What You Are Told: Visual Attention Sink in Large Multimodal Models
Seil Kang, Jinyeong Kim, Junhyeok Kim, Seong Jae Hwang
2025-03-06
Machine Learning · Computer Science
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
Yuval Ran-Milo, Hila Ofek, Shahar Mendel
2026-04-17
Machine Learning · Computer Science
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
Runyu Peng, Ruixiao Li, Mingshu Chen, Yunhua Zhou +2
2026-03-10