English

From $O(mn)$ to $O(r^2)$: Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency

Machine Learning 2026-02-10 v1 Artificial Intelligence

Abstract

As foundation models continue to scale, pretraining increasingly relies on data-parallel distributed optimization, making bandwidth-limited gradient synchronization a key bottleneck. Orthogonally, projection-based low-rank optimizers were mainly designed for memory efficiency, but remain suboptimal for communication-limited training: one-sided synchronization still transmits an O(rn)O(rn) object for an m×nm\times n matrix gradient and refresh steps can dominate peak communicated bytes. We propose TSR, which brings two-sided low-rank communication to Adam-family updates (TSR-Adam) by synchronizing a compact core UGVRr×rU^\top G V\in\mathbb{R}^{r\times r}, reducing the dominant per-step payload from O(mn)O(mn) to O(r2)O(r^2) while keeping moment states in low-dimensional cores. To further reduce the peak communication from subspace refresh, TSR-Adam adopts a randomized SVD-based refresh that avoids full-gradient synchronization. We additionally extend low-rank communication to embedding gradients with embedding-specific ranks and refresh schedules, yielding additional communication and memory savings over keeping embeddings dense. Across pretraining from 60M to 1B model scales, TSR-Adam reduces average communicated bytes per step by 13×13\times, and on GLUE fine-tuning it reduces communication by 25×25\times, while achieving comparable performance; we further provide a theoretical stationarity analysis for the proposed update. Code is available at https://github.com/DKmiyan/TSR-Adam.

Keywords

Cite

@article{arxiv.2602.08007,
  title  = {From $O(mn)$ to $O(r^2)$: Two-Sided Low-Rank Communication for Adam in Distributed Training with Memory Efficiency},
  author = {Sizhe Dang and Jiaqi Shao and Xiaodong Zheng and Guang Dai and Yan Song and Haishan Ye},
  journal= {arXiv preprint arXiv:2602.08007},
  year   = {2026}
}