English

Training Neural Networks at Any Scale

Machine Learning 2025-11-17 v1

Abstract

This article reviews modern optimization methods for training neural networks with an emphasis on efficiency and scale. We present state-of-the-art optimization algorithms under a unified algorithmic template that highlights the importance of adapting to the structures in the problem. We then cover how to make these algorithms agnostic to the scale of the problem. Our exposition is intended as an introduction for both practitioners and researchers who wish to be involved in these exciting new developments.

Keywords

Cite

@article{arxiv.2511.11163,
  title  = {Training Neural Networks at Any Scale},
  author = {Thomas Pethick and Kimon Antonakopoulos and Antonio Silveti-Falls and Leena Chennuru Vankadara and Volkan Cevher},
  journal= {arXiv preprint arXiv:2511.11163},
  year   = {2025}
}
R2 v1 2026-07-01T07:37:14.726Z