English

FLARE: Fast Low-rank Attention Routing Engine

Machine Learning 2026-02-03 v3

Abstract

The quadratic complexity of self-attention limits the scalability of transformers on long sequences. We introduce Fast Low-rank Attention Routing Engine (FLARE), a token-mixing operator that realizes low-rank attention by routing information through a small set of latent tokens. Each layer induces an input-input token mixing matrix of rank at most MM via a minimal encode-decode factorization implemented using only two standard scaled dot-product attention (SDPA) calls. Because the dominant O(NM){O}(NM) computation is expressed purely in terms of standard SDPA, FLARE is compatible with fused attention kernels and avoids materializing M×NM\times N projection matrices. FLARE further assigns disjoint latent slices to each attention head, yielding a mixture of head-specific low-rank pathways. Empirically, FLARE scales to one-million-point unstructured meshes on a single GPU, achieves state-of-the-art accuracy on PDE surrogate benchmarks, and outperforms general-purpose efficient-attention methods on the Long Range Arena suite. We additionally release a large-scale additive manufacturing benchmark dataset. Our code is available at https://github.com/vpuri3/FLARE.py.

Keywords

Cite

@article{arxiv.2508.12594,
  title  = {FLARE: Fast Low-rank Attention Routing Engine},
  author = {Vedant Puri and Aditya Joglekar and Sri Datta Ganesh Bandreddi and Kevin Ferguson and Yu-hsuan Chen and Yongjie Jessica Zhang and Levent Burak Kara},
  journal= {arXiv preprint arXiv:2508.12594},
  year   = {2026}
}
R2 v1 2026-07-01T04:54:10.278Z