English
Related papers

Related papers: Every Attention Matters: An Efficient Hybrid Archi…

200 papers

This work explores the challenge of building ``Machines that Can Remember'', framing long-term memory as the problem of efficient ultra-long context modeling. We argue that this requires three key properties: \textbf{sparsity},…

Computation and Language · Computer Science 2025-12-01 Xiang Hu , Zhanchao Zhou , Ruiqi Liang , Zehuan Li , Wei Wu , Jianguo Li

Hybrid attention architectures are becoming an increasingly important paradigm for improving LLM inference efficiency while preserving model quality, making hybrid architecture design a central problem. Existing designs often rely on manual…

Machine Learning · Computer Science 2026-05-21 Weizhe Chen , Miao Zhang , Junpeng Jiang , Yaping Li , Weili Guan , Liqiang Nie

Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of natural language processing tasks. These capabilities stem primarily from the self-attention mechanism, which enables modeling of long-range…

Computation and Language · Computer Science 2026-01-05 Zeng You , Yaofo Chen , Shuhai Zhang , Zhijie Qiu , Tingyu Wu , Yingjian Li , Yaowei Wang , Mingkui Tan

We develop hybrid memory architectures for general-purpose sequence processing neural networks, that combine key-value memory using softmax attention (KV-memory) with fast weight memory through dynamic synaptic modulation (FW-memory) -- the…

Machine Learning · Computer Science 2025-10-24 Kazuki Irie , Morris Yau , Samuel J. Gershman

This study introduces bifurcated attention, a method designed to enhance language model inference in shared-context batch decoding scenarios. Our approach addresses the challenge of redundant memory IO costs, a critical factor contributing…

We argue that neither transformers nor sub-quadratic architectures are well suited to training at long sequence lengths: the cost of processing the context is too expensive in the former, too inexpensive in the latter. Approaches such as…

Machine Learning · Computer Science 2025-07-08 Carles Gelada , Jacob Buckman , Sean Zhang , Txus Bach

Modeling long sequences of user behaviors has emerged as a critical frontier in generative recommendation. However, existing solutions face a dilemma: linear attention mechanisms achieve efficiency at the cost of retrieval precision due to…

Information Retrieval · Computer Science 2026-02-23 Lei Xin , Yuhao Zheng , Ke Cheng , Changjiang Jiang , Zifan Zhang , Fanhu Zeng

Sparse attention has been proposed as a way to alleviate the quadratic cost of transformers, a central bottleneck in long-context training. A promising line of work is $\alpha$-entmax attention, a differentiable sparse alternative to…

Machine Learning · Computer Science 2026-04-17 Nuno Gonçalves , Hugo Pitorro , Vlad Niculae , Edoardo Ponti , Lei Li , Andre Martins , Marcos Treviso

Mainstream Transformer-based large language models face major efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly, limiting long-context processing. Building large…

Attention mechanisms have become integral to modern convolutional neural networks (CNNs), delivering notable performance improvements with minimal computational overhead. However, the efficiency accuracy trade off of different channel…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Prem Babu Kanaparthi , Tulasi Venkata Sri Varshini Padamata

Long-context language modeling is commonly framed as a scalability challenge of token-level attention, yet local-to-global information structuring remains largely implicit in existing approaches. Drawing on cognitive theories of discourse…

Computation and Language · Computer Science 2026-04-10 Xiangyu Zeng , Qi Xu , Yunke Wang , Chang Xu

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B…

Artificial Intelligence · Computer Science 2026-05-27 MiniMax , : , Aili Chen , Aonian Li , Baichuan Zhou , Bangwei Gong , Binyang Jiang , Boji Dan , Changqing Yu , Chao Wang , Cheng Ma , Cheng Zhong , Cheng Zhu , Chengjun Xiao , Chengyi Yang , Chengyu Du , Chenyang Zhang , Chi Zhang , Chuangyi Huang , Chunhao Zhang , Chunhui Du , Chunyu Zhao , Congchao Guo , Da Chen , Deming Ding , Dianjun Sun , Dongyu Zhang , Enhui Yang , Fei Yu , Guang Zheng , Guodong Zheng , Guohong Li , Haichao Zhu , Haigang Zhou , Haimo Zhang , Han Ding , Hao Zhang , Haohai Sun , Haolin Lyu , Haonan Lu , Haoyu Wang , Huajie Shi , Huiyang Li , Jiacheng Chen , Jian Zhang , Jiaqi Zhuang , Jiaren Cai , Jiaxin Pan , Jiayao Li , Jiayuan Song , Jichuan Zhang , Jie Wang , Jihao Gu , Jin Zhu , Jingwei Dong , Jingyang Li , Jingyu Zhang , Jingze Zhuang , Jinhao Tian , Jinli Liu , Jinyi Hu , Jun Tao , Jun Zhang , Junbin Ruan , Junhao Xu , Junjie Yan , Junteng Liu , Junxian He , Kang Xu , Ke Ji , Ke Yang , Kecheng Xiao , Keyu Duan , Keyu Li , Le Han , Letian Ruan , Li Yuan , Lianfei Yu , Liheng Feng , Lijie Mo , Lin Li , Lingye Bao , Lingyu Yang , Lingyuan Zhou , Loki , Lu Chen , Lunbin Ceng , Ming Li , Ming Zhong , Mingliang Tao , Mingyuan Chi , Mujie Lin , Nan Hu , Ningxin Chen , Peiyin Zhu , Peng Gao , Pengcheng Gao , Pengfei Li , Penglin Li , Pengyu Zhao , Qibin Ren , Qidi Xu , Qihan Ren , Qile Li , Qin Wang , Quanliang Chen , Qunhong Ceng , Rong Tian , Rui Dong , Ruitao Leng , Ruize Zhang , Shanqi Liu , Shaoyu Chen , Sheng Jia , Shun Yao , Shuoran Zhao , Shuqi Yu , Sichen Li , Sicheng Pan , Songquan Zhu , Tengfei Li , Tian Xie , Tiancheng Qin , Tianrun Liang , Wei Liu , Weiqi Xu , Weitao Li , Weixiang Chen , Weiyu Cheng , Weiyu Zhang , Wenhu Chen , Wenqian Zhao , Xiancai Chen , Xiangjun Song , Xiangyuan Wang , Xiao Luo , Xiao Su , Xiaobo Li , Xiaodong Han , Xiaojie Wu , Xihao Song , Xingyi Han , Xinyu Guan , Xuan Lu , Xun Zou , Xunhao Lai , Xutong Li , Yan Gong , Yang Wang , Yang Xu , Yangsen Wang , Ye Tang , Yicheng Chen , Yinran Qiu , Yiqi Shi , Yiting Guo , Yiwen Huang , Yixuan Wang , Yongyi Hu , Yu Gao , Yu Zhang , Yuanxiang Ying , Yuanzhen Zhang , Yubo Wang , Yuchen Song , Yufeng Yang , Yuhang Meng , Yuhang Miao , Yuhao Li , Yujie Liu , Yulin Hu , Yunan Huang , Yunji Li , Yunyi Huang , Yusen Zhang , Yusu Hong , Yutao Xie , Yutong Zhang , Yuwen Liao , Yuxuan Shi , Yuze Wenren , Zebin Li , Zehan Li , Zejian Luo , Zeyu Jin , Zeyuan Sun , Zhanpeng Zhou , Zhaochen Su , Zhendong Li , Zhengmao Zhu , Zhengyuan Peng , Zhenhua Fan , Zhi Zhang , Zhichao Xu , Zhiheng Lv , Zhikang Xu , Zhitao He , Zhiwei He , Zhongyuan Li , Zibo Gao , Zijia Wu , Zijian Song , Zijian Zhou , Zijun Sun , Zishan Huang , Ziying Chen , Ziyue Ge

Standard softmax self-attention excels in vision tasks but incurs quadratic complexity O(N^2), limiting high-resolution deployment. Linear attention reduces the cost to O(N), yet its compressed state representations can impair modeling…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Ruibang Li , Guan Luo , Yiwei Zhang , Jin Gao , Bing Li , Weiming Hu

Interactive segmentation (IS) improves annotation efficiency by segmenting target regions from user prompts, with widespread applications in real-world scenarios. Current approaches face a critical trade-off: dense-token methods achieve…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 You Huang , Lichao Chen , Jiayi Ji , Liujuan Cao , Shengchuan Zhang , Rongrong Ji

In this paper, we propose a vision model that adopts token mixing, sequence-pooling, and convolutional tokenizers to achieve state-of-the-art performance and efficient inference in fixed context-length tasks. In the CIFAR100 benchmark, our…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Simpenzwe Honore Leandre , Natenaile Asmamaw Shiferaw , Dillip Rout

Large Language Models (LLMs) face significant challenges in long-context processing, including quadratic computational costs, information forgetting, and the context fragmentation inherent in retrieval-augmented generation (RAG). We propose…

Computation and Language · Computer Science 2026-02-10 Zhuoen Chen , Dongfang Li , Meishan Zhang , Baotian Hu , Min Zhang

Many advanced Large Language Model (LLM) applications require long-context processing, but the self-attention module becomes a bottleneck during the prefilling stage of inference due to its quadratic time complexity with respect to sequence…

Machine Learning · Computer Science 2025-06-02 Xiaodong Ji , Hailin Zhang , Fangcheng Fu , Bin Cui

Linear attentions have shown potential for improving Transformer efficiency, reducing attention's quadratic complexity to linear in sequence length. This holds exciting promise for (1) training linear Transformers from scratch, (2)…

Machine Learning · Computer Science 2024-02-08 Michael Zhang , Kush Bhatia , Hermann Kumbong , Christopher Ré

We study efficient reasoning under tight compute. We ask how to make structured, correct decisions without increasing test time cost. We add two training only components to small and medium Transformers that also transfer to broader…

Machine Learning · Computer Science 2026-03-11 Rian Atri

Transformer-based language models have recently been at the forefront of active research in text generation. However, these models' advances come at the price of prohibitive training costs, with parameter counts in the billions and compute…

Computation and Language · Computer Science 2025-02-04 Gabriel Lindenmaier , Sean Papay , Sebastian Padó