English

RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference

Computation and Language 2026-02-06 v1

Abstract

The quadratic complexity of attention mechanisms poses a critical bottleneck for large language models processing long contexts. While dynamic sparse attention methods offer input-adaptive efficiency, they face fundamental trade-offs: requiring preprocessing, lacking global evaluation, violating query independence, or incurring high computational overhead. We present RRAttention, a novel dynamic sparse attention method that simultaneously achieves all desirable properties through a head \underline{r}ound-\underline{r}obin (RR) sampling strategy. By rotating query sampling positions across attention heads within each stride, RRAttention maintains query independence while enabling efficient global pattern discovery with stride-level aggregation. Our method reduces complexity from O(L2)O(L^2) to O(L2/S2)O(L^2/S^2) and employs adaptive Top-τ\tau selection for optimal sparsity. Extensive experiments on natural language understanding (HELMET) and multimodal video comprehension (Video-MME) demonstrate that RRAttention recovers over 99\% of full attention performance while computing only half of the attention blocks, achieving 2.4×\times speedup at 128K context length and outperforming existing dynamic sparse attention methods.

Keywords

Cite

@article{arxiv.2602.05853,
  title  = {RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference},
  author = {Siran Liu and Guoxia Wang and Sa Wang and Jinle Zeng and HaoYang Xie and Siyu Lou and JiaBin Yang and DianHai Yu and Haifeng Wang and Chao Yang},
  journal= {arXiv preprint arXiv:2602.05853},
  year   = {2026}
}