Computation and Language · Computer Science
Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
Jiaqi Leng, Xiang Hu, Junxiong Wang, Jianguo Li +2
2026-05-01
Computation and Language · Computer Science
Zebra: Extending Context Window with Layerwise Grouped Local-Global Attention
Kaiqiang Song, Xiaoyang Wang, Sangwoo Cho, Xiaoman Pan +1
2023-12-15
Artificial Intelligence · Computer Science
AllMem: A Memory-centric Recipe for Efficient Long-context Modeling
Ziming Wang, Xiang Wang, Kailong Peng, Lang Qin +4
2026-02-17
Computation and Language · Computer Science
LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang +4
2024-07-12
Distributed, Parallel, and Cluster Computing · Computer Science
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Bin Lin, Chen Zhang, Tao Peng, Hanyu Zhao +11
2024-07-08
Computation and Language · Computer Science
Latent-Condensed Transformer for Efficient Long Context Modeling
Zeng You, Yaofo Chen, Qiuwu Chen, Ying Sun +4
2026-04-17
Computation and Language · Computer Science
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
Qianchao Zhu, Jiangfei Duan, Chang Chen, Siran Liu +5
2025-09-04
Artificial Intelligence · Computer Science
Solving Context Window Overflow in AI Agents
Anton Bulle Labate, Valesca Moura de Sousa, Sandro Rama Fiorini, Leonardo Guerreiro Azevedo +2
2025-12-01
Computation and Language · Computer Science
Exploring Context Window of Large Language Models via Decomposed Positional Vectors
Zican Dong, Junyi Li, Xin Men, Wayne Xin Zhao +4
2024-11-19
Computation and Language · Computer Science
Core Context Aware Transformers for Long Context Language Modeling
Yaofo Chen, Zeng You, Shuhai Zhang, Haokun Li +3
2025-08-05
Computation and Language · Computer Science
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
Sotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci +2
2024-06-03
Machine Learning · Computer Science
HyperAttention: Long-context Attention in Near-Linear Time
Insu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni +2
2023-12-04
Computation and Language · Computer Science
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
Hanzhi Zhang, Heng Fan, Kewei Sha, Yan Huang +1
2025-06-16
Computation and Language · Computer Science
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
MiniCPM Team, Wenhao An, Yingfa Chen, Yewei Fang +43
2026-03-03
Computation and Language · Computer Science
Training-Free Long-Context Scaling of Large Language Models
Chenxin An, Fei Huang, Jun Zhang, Shansan Gong +3
2024-05-30
Machine Learning · Computer Science
Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context Parallelism
Tao Bu, Qiangang Wang, Bowen Zeng, Hanwen Sun +3
2025-10-22
Computation and Language · Computer Science
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong +3
2026-02-03
Machine Learning · Computer Science
Megalodon: Efficient LLM Pretraining and Inference with Unlimited Context Length
Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen +6
2024-04-17
Computation and Language · Computer Science
LongHeads: Multi-Head Attention is Secretly a Long Context Processor
Yi Lu, Xin Zhou, Wei He, Jun Zhao +4
2024-03-26
Computation and Language · Computer Science
Training-free Context-adaptive Attention for Efficient Long Context Modeling
Zeng You, Yaofo Chen, Shuhai Zhang, Zhijie Qiu +4
2026-01-05