English
Related papers

Related papers: Efficient Context Scaling with LongCat ZigZag Atte…

200 papers

We present LongQLoRA, an efficient and effective method to extend context length of large language models with less training resources. LongQLoRA combines the advantages of Position Interpolation, QLoRA and Shift Short Attention of…

Computation and Language · Computer Science 2023-11-10 Jianxin Yang

Large Reasoning Models (LRMs) have shown promising accuracy improvements on complex problem-solving tasks. While these models have attained high accuracy by leveraging additional computation at test time, they need to generate long…

Computation and Language · Computer Science 2025-12-16 Coleman Hooper , Sebastian Zhao , Luca Manolache , Sehoon Kim , Michael W. Mahoney , Yakun Sophia Shao , Kurt Keutzer , Amir Gholami

Large language models have an exceptional capability to incorporate new information in a contextual manner. However, the full potential of such an approach is often restrained due to a limitation in the effective context length. One…

Computation and Language · Computer Science 2023-12-01 Szymon Tworkowski , Konrad Staniszewski , Mikołaj Pacek , Yuhuai Wu , Henryk Michalewski , Piotr Miłoś

Scaling inference for large language models (LLMs) is increasingly constrained by limited GPU memory, especially due to growing key-value (KV) caches required for long-context generation. While existing approaches offload KV caches to CPU…

Machine Learning · Computer Science 2025-07-08 Weishu Deng , Yujie Yang , Peiran Du , Lingfeng Xiang , Zhen Lin , Chen Zhong , Song Jiang , Hui Lu , Jia Rao

The growing demand for long-context inference capabilities in Large Language Models (LLMs) has intensified the computational and memory bottlenecks inherent to the self-attention mechanism. To address this challenge, we introduce BLASST, a…

We introduce LongCat-Flash, a 560-billion-parameter Mixture-of-Experts (MoE) language model designed for both computational efficiency and advanced agentic capabilities. Stemming from the need for scalable efficiency, LongCat-Flash adopts…

Computation and Language · Computer Science 2025-09-22 Meituan LongCat Team , Bayan , Bei Li , Bingye Lei , Bo Wang , Bolin Rong , Chao Wang , Chao Zhang , Chen Gao , Chen Zhang , Cheng Sun , Chengcheng Han , Chenguang Xi , Chi Zhang , Chong Peng , Chuan Qin , Chuyu Zhang , Cong Chen , Congkui Wang , Dan Ma , Daoru Pan , Defei Bu , Dengchang Zhao , Deyang Kong , Dishan Liu , Feiye Huo , Fengcun Li , Fubao Zhang , Gan Dong , Gang Liu , Gang Xu , Ge Li , Guoqiang Tan , Guoyuan Lin , Haihang Jing , Haomin Fu , Haonan Yan , Haoxing Wen , Haozhe Zhao , Hong Liu , Hongmei Shi , Hongyan Hao , Hongyin Tang , Huantian Lv , Hui Su , Jiacheng Li , Jiahao Liu , Jiahuan Li , Jiajun Yang , Jiaming Wang , Jian Yang , Jianchao Tan , Jiaqi Sun , Jiaqi Zhang , Jiawei Fu , Jiawei Yang , Jiaxi Hu , Jiayu Qin , Jingang Wang , Jiyuan He , Jun Kuang , Junhui Mei , Kai Liang , Ke He , Kefeng Zhang , Keheng Wang , Keqing He , Liang Gao , Liang Shi , Lianhui Ma , Lin Qiu , Lingbin Kong , Lingtong Si , Linkun Lyu , Linsen Guo , Liqi Yang , Lizhi Yan , Mai Xia , Man Gao , Manyuan Zhang , Meng Zhou , Mengxia Shen , Mingxiang Tuo , Mingyang Zhu , Peiguang Li , Peng Pei , Peng Zhao , Pengcheng Jia , Pingwei Sun , Qi Gu , Qianyun Li , Qingyuan Li , Qiong Huang , Qiyuan Duan , Ran Meng , Rongxiang Weng , Ruichen Shao , Rumei Li , Shizhe Wu , Shuai Liang , Shuo Wang , Suogui Dang , Tao Fang , Tao Li , Tefeng Chen , Tianhao Bai , Tianhao Zhou , Tingwen Xie , Wei He , Wei Huang , Wei Liu , Wei Shi , Wei Wang , Wei Wu , Weikang Zhao , Wen Zan , Wenjie Shi , Xi Nan , Xi Su , Xiang Li , Xiang Mei , Xiangyang Ji , Xiangyu Xi , Xiangzhou Huang , Xianpeng Li , Xiao Fu , Xiao Liu , Xiao Wei , Xiaodong Cai , Xiaolong Chen , Xiaoqing Liu , Xiaotong Li , Xiaowei Shi , Xiaoyu Li , Xili Wang , Xin Chen , Xing Hu , Xingyu Miao , Xinyan He , Xuemiao Zhang , Xueyuan Hao , Xuezhi Cao , Xunliang Cai , Xurui Yang , Yan Feng , Yang Bai , Yang Chen , Yang Yang , Yaqi Huo , Yerui Sun , Yifan Lu , Yifan Zhang , Yipeng Zang , Yitao Zhai , Yiyang Li , Yongjing Yin , Yongkang Lv , Yongwei Zhou , Yu Yang , Yuchen Xie , Yueqing Sun , Yuewen Zheng , Yuhuai Wei , Yulei Qian , Yunfan Liang , Yunfang Tai , Yunke Zhao , Zeyang Yu , Zhao Zhang , Zhaohua Yang , Zhenchao Zhang , Zhikang Xia , Zhiye Zou , Zhizhao Zeng , Zhongda Su , Zhuofan Chen , Zijian Zhang , Ziwen Wang , Zixu Jiang , Zizhe Zhao , Zongyu Wang , Zunhai Su

Self-attention and position embedding are two key modules in transformer-based Large Language Models (LLMs). However, the potential relationship between them is far from well studied, especially for long context window extending. In fact,…

Machine Learning · Computer Science 2024-02-29 Shiyi Zhu , Jing Ye , Wei Jiang , Siqiao Xue , Qi Zhang , Yifan Wu , Jianguo Li

Training and serving long-context large language models (LLMs) incurs substantial overhead. To address this, two critical steps are often required: a pretrained LLM typically undergoes a separate stage for context length extension by…

Computation and Language · Computer Science 2024-12-06 Suyu Ge , Xihui Lin , Yunan Zhang , Jiawei Han , Hao Peng

Large Language Models (LLMs) have demonstrated remarkable capabilities across various applications, but their performance on long-context tasks is often limited by the computational complexity of attention mechanisms. We introduce a novel…

Machine Learning · Computer Science 2025-02-25 Bo Chen , Yingyu Liang , Zhizhou Sha , Zhenmei Shi , Zhao Song

Long context understanding remains challenging for large language models due to their limited context windows. This paper introduces Long Input Fine-Tuning (LIFT) for long context modeling, a novel framework that enhances LLM performance on…

Computation and Language · Computer Science 2024-12-19 Yansheng Mao , Jiaqi Li , Fanxu Meng , Jing Xiong , Zilong Zheng , Muhan Zhang

Large language models (LLMs) now support extremely long context windows, but the quadratic complexity of vanilla attention results in significantly long Time-to-First-Token (TTFT) latency. Existing approaches to address this complexity…

Computation and Language · Computer Science 2025-09-04 Qianchao Zhu , Jiangfei Duan , Chang Chen , Siran Liu , Guanyu Feng , Xin Lv , Xiao Chuanfu , Dahua Lin , Chao Yang

Recent advances in transformer-based Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their quadratic computational complexity concerning sequence length remains a significant bottleneck…

Computation and Language · Computer Science 2025-06-05 Zichuan Fu , Wentao Song , Yejing Wang , Xian Wu , Yefeng Zheng , Yingying Zhang , Derong Xu , Xuetao Wei , Tong Xu , Xiangyu Zhao

Transformer-based Large Language Models (LLMs) have exhibited remarkable success in extensive tasks primarily attributed to self-attention mechanism, which requires a token to consider all preceding tokens as its context to compute…

Computation and Language · Computer Science 2025-08-05 Yaofo Chen , Zeng You , Shuhai Zhang , Haokun Li , Yirui Li , Yaowei Wang , Mingkui Tan

The computational burden of attention in long-context language models has motivated two largely independent lines of work: sparse attention mechanisms that reduce complexity by attending to selected tokens, and gated attention variants that…

Artificial Intelligence · Computer Science 2026-01-23 Alfred Shen , Aaron Shen

Long-context models are essential for many applications but face inefficiencies in loading large KV caches during decoding. Prior methods enforce fixed token budgets for sparse attention, assuming a set number of tokens can approximate full…

Machine Learning · Computer Science 2025-02-19 Kan Zhu , Tian Tang , Qinyu Xu , Yile Gu , Zhichen Zeng , Rohan Kadekodi , Liangyu Zhao , Ang Li , Arvind Krishnamurthy , Baris Kasikci

Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking, each target token's effective context remains short. We introduce EXACT, a…

Computation and Language · Computer Science 2026-05-12 Jinchang Zhu , Jindong Li , Chengyu Zou , Rong Fu , Chao Wang , Haowei He , Menglin Yang

Large language models (LLMs) now support context windows of hundreds of thousands to millions of tokens, enabling applications such as long-document summarization, large-scale code synthesis, multi-document question answering and persistent…

Computation and Language · Computer Science 2025-10-22 Siyuan Yan , Guo-Qing Jiang , Yuchen Zhang , Xiaoxing Ma , Ran Zhu , Chun Cao , Jingwei Xu

Scaling language models to handle longer contexts introduces substantial memory challenges due to the growing cost of key-value (KV) caches. Motivated by the efficiency gains of hybrid models and the broad availability of pretrained large…

Computation and Language · Computer Science 2026-05-19 Xuan Zhang , Fengzhuo Zhang , Cunxiao Du , Chao Du , Tianyu Pang , Wei Gao , Min Lin

We present core attention disaggregation (CAD), a technique that improves long-context large language model training by decoupling the core attention computation, softmax(QK^T)V, from the rest of the model and executing it on a separate…

Machine Learning · Computer Science 2025-10-22 Yonghao Zhuang , Junda Chen , Bo Pang , Yi Gu , Yibo Zhu , Yimin Jiang , Ion Stoica , Eric Xing , Hao Zhang

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. While various sparse attention…

Computation and Language · Computer Science 2026-03-09 Qihang Fan , Huaibo Huang , Zhiying Wu , Juqiu Wang , Bingning Wang , Ran He