English
Related papers

Related papers: The Structural Origin of Attention Sink: Variance …

200 papers

Large Language Models (LLMs) often allocate disproportionate attention to specific tokens, a phenomenon commonly referred to as the attention sink. While such sinks are generally considered detrimental, prior studies have identified a…

Machine Learning · Computer Science 2026-03-10 Runyu Peng , Ruixiao Li , Mingshu Chen , Yunhua Zhou , Qipeng Guo , Xipeng Qiu

Language Models (LMs) assign significant attention to the first token, even if it is not semantically important, which is known as attention sink. This phenomenon has been widely adopted in applications such as streaming/long context…

Computation and Language · Computer Science 2025-03-04 Xiangming Gu , Tianyu Pang , Chao Du , Qian Liu , Fengzhuo Zhang , Cunxiao Du , Ye Wang , Min Lin

Large Language Models (LLMs) often assign disproportionate attention to the first token, a phenomenon known as the attention sink. Several recent approaches aim to address this issue, including Sink Attention in GPT-OSS and Gated Attention…

Computation and Language · Computer Science 2026-05-28 Zizhuo Fu , Wenxuan Zeng , Runsheng Wang , Meng Li

Large Language Models (LLMs), despite their impressive capabilities, often fail to accurately repeat a single word when prompted to, and instead output unrelated text. This unexplained failure mode represents a vulnerability, allowing even…

Machine Learning · Computer Science 2025-03-13 Itay Yona , Ilia Shumailov , Jamie Hayes , Federico Barbero , Yossi Gandelsman

We study two recurring phenomena in Transformer language models: massive activations, in which a small number of tokens exhibit extreme outliers in a few channels, and attention sinks, in which certain tokens attract disproportionate…

Artificial Intelligence · Computer Science 2026-03-06 Shangwen Sun , Alfredo Canziani , Yann LeCun , Jiachen Zhu

Transformers commonly exhibit an attention sink: disproportionately high attention to the first position. We study this behavior in GPT-2-style models with learned query biases and absolute positional embeddings. Combining structural…

Machine Learning · Computer Science 2026-04-17 Yuval Ran-Milo , Hila Ofek , Shahar Mendel

Large Language Models (LLMs) tend to attend heavily to the first token in the sequence -- creating a so-called attention sink. Many works have studied this phenomenon in detail, proposing various ways to either leverage or alleviate it.…

Large language models (LLMs) often concentrate their attention on a few specific tokens referred to as attention sinks. Common examples include the first token, a prompt-independent sink, and punctuation tokens, which are prompt-dependent.…

Computation and Language · Computer Science 2025-09-23 Stephen Zhang , Mustafa Khan , Vardan Papyan

Attention sinks and massive activations are recurring and closely related phenomena in Transformer models. Existing explanations have largely focused on the forward pass, yet in pre-norm Transformers, large residual-stream norms play only…

Machine Learning · Computer Science 2026-05-07 Yihong Chen , Zhouchen Lin , Quanming Yao

Attention sinks are tokens, often the beginning-of-sequence (BOS) token, that receive disproportionately high attention despite limited semantic relevance. In this work, we identify a class of attention sinks, which we term secondary sinks,…

Machine Learning · Computer Science 2026-03-17 Jeffrey T. H. Wong , Cheng Zhang , Louis Mahon , Wayne Luk , Anton Isopoussu , Yiren Zhao

Transformers underpin modern large language models (LLMs) and are commonly assumed to be behaviorally unstructured at random initialization, with all meaningful preferences emerging only through large-scale training. We challenge this…

Machine Learning · Statistics 2026-02-06 Siquan Li , Yao Tong , Haonan Wang , Tianyang Hu

The goal of this paper is to strengthen the reasoning of Omnimodal Large Language Models (Omni-LLMs) at inference time, without additional training. These models jointly process video, audio, and text, and given the large number of tokens…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Suho Yoo , Youngjoon Jang , Joon Son Chung

Large language models (LLMs) have recently advanced auditory speech recognition (ASR), visual speech recognition (VSR), and audio-visual speech recognition (AVSR). However, understanding of their internal dynamics under fine-tuning remains…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-28 Anand , Umberto Cappellazzo , Stavros Petridis , Maja Pantic

Vision transformers have emerged as a powerful tool across a wide range of applications, yet their inner workings remain only partially understood. In this work, we examine the phenomenon of massive tokens - tokens with exceptionally high…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Andrew Lu , Wentinn Liao , Liuhui Wang , Huzheng Yang , Jianbo Shi

Attention sinks -- tokens that receive disproportionate attention mass -- are assumed to be functionally important in autoregressive language models, but their role in diffusion transformers remains unclear. We present a causal analysis in…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Fangzheng Wu , Brian Summa

Large language models frequently exhibit hallucinations: fluent and confident outputs that are factually incorrect or unsupported by the input context. While recent hallucination detection methods have explored various features derived from…

Computation and Language · Computer Science 2026-04-14 Jakub Binkowski , Kamil Adamczewski , Tomasz Kajdanowicz

Practitioners have consistently observed three puzzling phenomena in transformer-based large language models (LLMs): attention sinks, value-state drains, and residual-state peaks, collectively referred to as extreme-token phenomena. These…

Machine Learning · Computer Science 2024-11-08 Tianyu Guo , Druv Pai , Yu Bai , Jiantao Jiao , Michael I. Jordan , Song Mei

The behavior of the network and its stability are governed by both dynamics of individual nodes as well as their topological interconnections. Attention mechanism as an integral part of neural network models was initially designed for…

Machine Learning · Computer Science 2022-12-20 Nooshin Bahador , Milad Lankarany

Large language models (LLMs) suffer from hallucination and context forgetting. Prior studies suggest that attention drift is a primary cause of these problems, where LLMs' focus shifts towards newly generated tokens and away from the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Xu Liu , Guikun Chen , Wenguan Wang

Harmful fine-tuning can invalidate safety alignment of large language models, exposing significant safety risks. In this paper, we utilize the attention sink mechanism to mitigate harmful fine-tuning. Specifically, we first measure a…

Artificial Intelligence · Computer Science 2026-02-12 Guozhi Liu , Weiwei Lin , Tiansheng Huang , Ruichao Mo , Qi Mu , Xiumin Wang , Li Shen
‹ Prev 1 2 3 10 Next ›