English

Sketching as a Tool for Understanding and Accelerating Self-attention for Long Sequences

Machine Learning 2021-12-13 v1 Computation and Language Machine Learning

Abstract

Transformer-based models are not efficient in processing long sequences due to the quadratic space and time complexity of the self-attention modules. To address this limitation, Linformer and Informer are proposed to reduce the quadratic complexity to linear (modulo logarithmic factors) via low-dimensional projection and row selection respectively. These two models are intrinsically connected, and to understand their connection, we introduce a theoretical framework of matrix sketching. Based on the theoretical analysis, we propose Skeinformer to accelerate self-attention and further improve the accuracy of matrix approximation to self-attention with three carefully designed components: column sampling, adaptive row normalization and pilot sampling reutilization. Experiments on the Long Range Arena (LRA) benchmark demonstrate that our methods outperform alternatives with a consistently smaller time/space footprint.

Keywords

Cite

@article{arxiv.2112.05359,
  title  = {Sketching as a Tool for Understanding and Accelerating Self-attention for Long Sequences},
  author = {Yifan Chen and Qi Zeng and Dilek Hakkani-Tur and Di Jin and Heng Ji and Yun Yang},
  journal= {arXiv preprint arXiv:2112.05359},
  year   = {2021}
}
R2 v1 2026-06-24T08:11:51.992Z