English

Minibatch and Local SGD: Algorithmic Stability and Linear Speedup in Generalization

Machine Learning 2025-10-14 v3 Artificial Intelligence

Abstract

The increasing scale of data propels the popularity of leveraging parallelism to speed up the optimization. Minibatch stochastic gradient descent (minibatch SGD) and local SGD are two popular methods for parallel optimization. The existing theoretical studies show a linear speedup of these methods with respect to the number of machines, which, however, is measured by optimization errors in a multi-pass setting. As a comparison, the stability and generalization of these methods are much less studied. In this paper, we study the stability and generalization analysis of minibatch and local SGD to understand their learnability by introducing an expectation-variance decomposition. We incorporate training errors into the stability analysis, which shows how small training errors help generalization for overparameterized models. We show minibatch and local SGD achieve a linear speedup to attain the optimal risk bounds.

Keywords

Cite

@article{arxiv.2310.01139,
  title  = {Minibatch and Local SGD: Algorithmic Stability and Linear Speedup in Generalization},
  author = {Yunwen Lei and Tao Sun and Mingrui Liu},
  journal= {arXiv preprint arXiv:2310.01139},
  year   = {2025}
}

Comments

Published in Applied and Computational Harmonic Analysis, October 2025

R2 v1 2026-06-28T12:38:12.821Z