English
Related papers

Related papers: Gated Sparse Attention: Combining Computational Ef…

200 papers

Sparse attention, which selectively attends to a subset of tokens in the context was supposed to be efficient. However, its theoretical reduction in FLOPs has rarely translated into wall-clock speed-up over its dense attention counterparts…

Computation and Language · Computer Science 2025-02-06 Xihui Lin , Yunan Zhang , Suyu Ge , Liliang Ren , Barun Patra , Vishrav Chaudhary , Hao Peng , Xia Song

Structured dilated attention has an appealing inference-time efficiency knob: it reduces the FLOPs of attention and the KV cache size by a factor of the dilation size D, while preserving long-range connectivity. While prior work studies it…

Machine Learning · Computer Science 2026-05-29 Xiuying Wei , Caglar Gulcehre

Reward sparsity in long-horizon reinforcement learning (RL) tasks remains a significant challenge, while existing outcome-based reward shaping struggles to define meaningful immediate rewards without introducing bias or requiring explicit…

Machine Learning · Computer Science 2025-08-15 Zetian Sun , Dongfang Li , Zhuoen Chen , Yuhuai Qin , Baotian Hu

Large language models (LLMs) increasingly require mechanisms for continual adaptation without full retraining. However, sequential updates can lead to catastrophic forgetting, where new edits degrade previously acquired knowledge. This work…

Machine Learning · Computer Science 2025-10-21 William Hoy , Nurcin Celik

We propose Sparse Sinkhorn Attention, a new efficient and sparse method for learning to attend. Our method is based on differentiable sorting of internal representations. Concretely, we introduce a meta sorting network that learns to…

Machine Learning · Computer Science 2020-02-27 Yi Tay , Dara Bahri , Liu Yang , Donald Metzler , Da-Cheng Juan

Transformer has achieved remarkable success in language, image, and speech processing. Recently, various efficient attention architectures have been proposed to improve transformer's efficiency while largely preserving its efficacy,…

Machine Learning · Computer Science 2025-01-14 Jun Zhang , Shuyang Jiang , Jiangtao Feng , Lin Zheng , Lingpeng Kong

Transformer models are computationally costly on long sequences since regular attention has quadratic $O(n^2)$ time complexity. We introduce Wavelet-Enhanced Random Spectral Attention (WERSA), a novel mechanism of linear $O(n)$ time…

Machine Learning · Computer Science 2025-07-14 Vincenzo Dentamaro

Large models based on the Transformer architecture are susceptible to extreme-token phenomena, such as attention sinks and value-state drains. These issues, which degrade model performance, quantization fidelity, and interpretability, arise…

Machine Learning · Computer Science 2026-01-27 Rui Bu , Haofeng Zhong , Wenzheng Chen , Yangyan Li

Sparse Attention is a technique that approximates standard attention computation with sub-quadratic complexity. This is achieved by selectively ignoring smaller entries in the attention matrix during the softmax function computation.…

Machine Learning · Computer Science 2025-02-13 Yichuan Deng , Zhao Song , Jing Xiong , Chiwun Yang

Improving the effectiveness and efficiency of large language models (LLMs) simultaneously is a critical yet challenging research goal. In this paper, we find that low-rank pre-training, normally considered as efficient methods that will…

Computation and Language · Computer Science 2024-11-05 Xingtai Lv , Ning Ding , Kaiyan Zhang , Ermo Hua , Ganqu Cui , Bowen Zhou

Large language models (LLMs) now support context windows of hundreds of thousands to millions of tokens, enabling applications such as long-document summarization, large-scale code synthesis, multi-document question answering and persistent…

Computation and Language · Computer Science 2025-10-22 Siyuan Yan , Guo-Qing Jiang , Yuchen Zhang , Xiaoxing Ma , Ran Zhu , Chun Cao , Jingwei Xu

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. The key technical breakthroughs of DeepSeek-V3.2 are as follows: (1) DeepSeek Sparse Attention (DSA): We…

Computation and Language · Computer Science 2025-12-03 DeepSeek-AI , Aixin Liu , Aoxue Mei , Bangcai Lin , Bing Xue , Bingxuan Wang , Bingzheng Xu , Bochao Wu , Bowei Zhang , Chaofan Lin , Chen Dong , Chengda Lu , Chenggang Zhao , Chengqi Deng , Chenhao Xu , Chong Ruan , Damai Dai , Daya Guo , Dejian Yang , Deli Chen , Erhang Li , Fangqi Zhou , Fangyun Lin , Fucong Dai , Guangbo Hao , Guanting Chen , Guowei Li , H. Zhang , Hanwei Xu , Hao Li , Haofen Liang , Haoran Wei , Haowei Zhang , Haowen Luo , Haozhe Ji , Honghui Ding , Hongxuan Tang , Huanqi Cao , Huazuo Gao , Hui Qu , Hui Zeng , Jialiang Huang , Jiashi Li , Jiaxin Xu , Jiewen Hu , Jingchang Chen , Jingting Xiang , Jingyang Yuan , Jingyuan Cheng , Jinhua Zhu , Jun Ran , Junguang Jiang , Junjie Qiu , Junlong Li , Junxiao Song , Kai Dong , Kaige Gao , Kang Guan , Kexin Huang , Kexing Zhou , Kezhao Huang , Kuai Yu , Lean Wang , Lecong Zhang , Lei Wang , Liang Zhao , Liangsheng Yin , Lihua Guo , Lingxiao Luo , Linwang Ma , Litong Wang , Liyue Zhang , M. S. Di , M. Y Xu , Mingchuan Zhang , Minghua Zhang , Minghui Tang , Mingxu Zhou , Panpan Huang , Peixin Cong , Peiyi Wang , Qiancheng Wang , Qihao Zhu , Qingyang Li , Qinyu Chen , Qiushi Du , Ruiling Xu , Ruiqi Ge , Ruisong Zhang , Ruizhe Pan , Runji Wang , Runqiu Yin , Runxin Xu , Ruomeng Shen , Ruoyu Zhang , S. H. Liu , Shanghao Lu , Shangyan Zhou , Shanhuang Chen , Shaofei Cai , Shaoyuan Chen , Shengding Hu , Shengyu Liu , Shiqiang Hu , Shirong Ma , Shiyu Wang , Shuiping Yu , Shunfeng Zhou , Shuting Pan , Songyang Zhou , Tao Ni , Tao Yun , Tian Pei , Tian Ye , Tianyuan Yue , Wangding Zeng , Wen Liu , Wenfeng Liang , Wenjie Pang , Wenjing Luo , Wenjun Gao , Wentao Zhang , Xi Gao , Xiangwen Wang , Xiao Bi , Xiaodong Liu , Xiaohan Wang , Xiaokang Chen , Xiaokang Zhang , Xiaotao Nie , Xin Cheng , Xin Liu , Xin Xie , Xingchao Liu , Xingkai Yu , Xingyou Li , Xinyu Yang , Xinyuan Li , Xu Chen , Xuecheng Su , Xuehai Pan , Xuheng Lin , Xuwei Fu , Y. Q. Wang , Yang Zhang , Yanhong Xu , Yanru Ma , Yao Li , Yao Li , Yao Zhao , Yaofeng Sun , Yaohui Wang , Yi Qian , Yi Yu , Yichao Zhang , Yifan Ding , Yifan Shi , Yiliang Xiong , Ying He , Ying Zhou , Yinmin Zhong , Yishi Piao , Yisong Wang , Yixiao Chen , Yixuan Tan , Yixuan Wei , Yiyang Ma , Yiyuan Liu , Yonglun Yang , Yongqiang Guo , Yongtong Wu , Yu Wu , Yuan Cheng , Yuan Ou , Yuanfan Xu , Yuduan Wang , Yue Gong , Yuhan Wu , Yuheng Zou , Yukun Li , Yunfan Xiong , Yuxiang Luo , Yuxiang You , Yuxuan Liu , Yuyang Zhou , Z. F. Wu , Z. Z. Ren , Zehua Zhao , Zehui Ren , Zhangli Sha , Zhe Fu , Zhean Xu , Zhenda Xie , Zhengyan Zhang , Zhewen Hao , Zhibin Gou , Zhicheng Ma , Zhigang Yan , Zhihong Shao , Zhixian Huang , Zhiyu Wu , Zhuoshu Li , Zhuping Zhang , Zian Xu , Zihao Wang , Zihui Gu , Zijia Zhu , Zilin Li , Zipeng Zhang , Ziwei Xie , Ziyi Gao , Zizheng Pan , Zongqing Yao , Bei Feng , Hui Li , J. L. Cai , Jiaqi Ni , Lei Xu , Meng Li , Ning Tian , R. J. Chen , R. L. Jin , S. S. Li , Shuang Zhou , Tianyu Sun , X. Q. Li , Xiangyue Jin , Xiaojin Shen , Xiaosha Chen , Xinnan Song , Xinyi Zhou , Y. X. Zhu , Yanping Huang , Yaohui Li , Yi Zheng , Yuchen Zhu , Yunxian Ma , Zhen Huang , Zhipeng Xu , Zhongyu Zhang , Dongjie Ji , Jian Liang , Jianzhong Guo , Jin Chen , Leyi Xia , Miaojun Wang , Mingming Li , Peng Zhang , Ruyi Chen , Shangmian Sun , Shaoqing Wu , Shengfeng Ye , T. Wang , W. L. Xiao , Wei An , Xianzu Wang , Xiaowen Sun , Xiaoxiang Wang , Ying Tang , Yukun Zha , Zekai Zhang , Zhe Ju , Zhen Zhang , Zihua Qu

Self-attention scales quadratically with input size, limiting its use for large-scale physical systems. Although sparse attention mechanisms provide a viable alternative, they are primarily designed for regular structures such as text or…

Machine Learning · Computer Science 2025-06-17 Catalin E. Brita , Hieu Nguyen , Lohithsai Yadala Chanchu , Domonkos Nagy , Maksim Zhdanov

Sparse autoencoders (SAEs) have recently emerged as a powerful tool for language model steering. Prior work has explored top-k SAE latents for steering, but we observe that many dimensions among the top-k latents capture non-semantic…

Computation and Language · Computer Science 2025-10-03 Jiaqing Xie

Transformers have demonstrated strong performance across a wide range of sequence modeling tasks, but their quadratic attention complexity limits scalability to long sequences. Linear models such as Mamba and sliding-window attention (SWA)…

Machine Learning · Computer Science 2025-09-03 Aref Jafari , Yuhe Fan , Benyamin Jamialahmadi , Parsa Farinneya , Boxing Chen , Marzieh S. Tahaei

Large language models with long context windows can answer complex questions directly from full-length academic, technical, and policy documents, but passing entire documents is often costly, slow, and can degrade answer quality while…

Sparse attention offers a promising strategy to extend long-context capabilities in Transformer LLMs, yet its efficiency-accuracy trade-offs remain unclear due to the lack of comprehensive evaluation. We address this gap with the…

Computation and Language · Computer Science 2026-01-28 Piotr Nawrot , Robert Li , Renjie Huang , Sebastian Ruder , Kelly Marchisio , Edoardo M. Ponti

Softmax-based dot-product attention is a cornerstone of Transformer architectures, enabling remarkable capabilities such as in-context learning. However, as context lengths increase, a fundamental limitation of the softmax function emerges:…

Machine Learning · Computer Science 2026-02-12 Sai Surya Duvvuri , Nirmal Patel , Nilesh Gupta , Inderjit S. Dhillon

As large language models (LLMs) and visual language models (VLMs) grow in scale and application, attention mechanisms have become a central computational bottleneck due to their high memory and time complexity. While many efficient…

Machine Learning · Computer Science 2025-07-11 Zhengyu Tian , Anantha Padmanaban Krishna Kumar , Hemant Krishnakumar , Reza Rawassizadeh

Self-attention mechanisms have achieved great success on a variety of NLP tasks due to its flexibility of capturing dependency between arbitrary positions in a sequence. For problems such as query-based summarization (Qsumm) and knowledge…

Computation and Language · Computer Science 2020-02-19 Yujia Xie , Tianyi Zhou , Yi Mao , Weizhu Chen