English

Scaffold with Stochastic Gradients: New Analysis with Linear Speed-Up

Machine Learning 2025-03-11 v1 Machine Learning Optimization and Control

Abstract

This paper proposes a novel analysis for the Scaffold algorithm, a popular method for dealing with data heterogeneity in federated learning. While its convergence in deterministic settings--where local control variates mitigate client drift--is well established, the impact of stochastic gradient updates on its performance is less understood. To address this problem, we first show that its global parameters and control variates define a Markov chain that converges to a stationary distribution in the Wasserstein distance. Leveraging this result, we prove that Scaffold achieves linear speed-up in the number of clients up to higher-order terms in the step size. Nevertheless, our analysis reveals that Scaffold retains a higher-order bias, similar to FedAvg, that does not decrease as the number of clients increases. This highlights opportunities for developing improved stochastic federated learning algorithms

Keywords

Cite

@article{arxiv.2503.07594,
  title  = {Scaffold with Stochastic Gradients: New Analysis with Linear Speed-Up},
  author = {Paul Mangold and Alain Durmus and Aymeric Dieuleveut and Eric Moulines},
  journal= {arXiv preprint arXiv:2503.07594},
  year   = {2025}
}
R2 v1 2026-06-28T22:14:28.907Z