Computation and Language · Computer Science
ReAttention: Training-Free Infinite Context with Finite Attention Scope
Xiaoran Liu, Ruixiao Li, Qipeng Guo, Zhigeng Liu +6
2025-03-20
Distributed, Parallel, and Cluster Computing · Computer Science
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Bin Lin, Chen Zhang, Tao Peng, Hanyu Zhao +11
2024-07-08
Machine Learning · Computer Science
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
Ruizhe Huang, Kexuan Zhang, Yihao Fang, Baifeng Yu
2026-01-01
Computation and Language · Computer Science
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
Chaojun Xiao, Pengle Zhang, Xu Han, Guangxuan Xiao +4
2024-05-29
Computer Vision and Pattern Recognition · Computer Science
Long-Short Transformer: Efficient Transformers for Language and Vision
Chen Zhu, Wei Ping, Chaowei Xiao, Mohammad Shoeybi +3
2021-12-08
Computation and Language · Computer Science
In-Context Former: Lightning-fast Compressing Context for Large Language Model
Xiangfeng Wang, Zaiyi Chen, Zheyong Xie, Tong Xu +2
2024-11-06
Computation and Language · Computer Science
Core Context Aware Transformers for Long Context Language Modeling
Yaofo Chen, Zeng You, Shuhai Zhang, Haokun Li +3
2025-08-05
Computation and Language · Computer Science
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
Qianchao Zhu, Jiangfei Duan, Chang Chen, Siran Liu +5
2025-09-04
Computation and Language · Computer Science
DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration
Hanzhi Zhang, Heng Fan, Kewei Sha, Yan Huang +1
2025-06-16
Computation and Language · Computer Science
Latent-Condensed Transformer for Efficient Long Context Modeling
Zeng You, Yaofo Chen, Qiuwu Chen, Ying Sun +4
2026-04-17
Computation and Language · Computer Science
Ltri-LLM: Streaming Long Context Inference for LLMs with Training-Free Dynamic Triangular Attention Pattern
Hongyin Tang, Di Xiu, Lanrui Wang, Xiurui Geng +2
2024-12-09
Computation and Language · Computer Science
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
MiniCPM Team, Wenhao An, Yingfa Chen, Yewei Fang +43
2026-03-03
Machine Learning · Computer Science
HyperAttention: Long-context Attention in Near-Linear Time
Insu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni +2
2023-12-04
Computation and Language · Computer Science
Context Memorization for Efficient Long Context Generation
Yasuyuki Okoshi, Hao Mark Chen, Guanxi Lu, Hongxiang Fan +2
2026-05-19
Computation and Language · Computer Science
LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models
Chi Han, Qifan Wang, Hao Peng, Wenhan Xiong +3
2024-06-26
Computation and Language · Computer Science
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
James Begin, Namit Agrawal, Eshan Singh, Yicheng Fu +3
2025-03-03
Computation and Language · Computer Science
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong +3
2026-02-03
Computation and Language · Computer Science
Training-free Context-adaptive Attention for Efficient Long Context Modeling
Zeng You, Yaofo Chen, Shuhai Zhang, Zhijie Qiu +4
2026-01-05