English

The Effectiveness of Local Updates for Decentralized Learning under Data Heterogeneity

Machine Learning 2024-12-25 v3 Optimization and Control

Abstract

We revisit two fundamental decentralized optimization methods, Decentralized Gradient Tracking (DGT) and Decentralized Gradient Descent (DGD), with multiple local updates. We consider two settings and demonstrate that incorporating local update steps can reduce communication complexity. Specifically, for μ\mu-strongly convex and LL-smooth loss functions, we proved that local DGT achieves communication complexity {}{O~(Lμ(K+1)+δ+μμ(1ρ)+ρ(1ρ)2L+δμ)\tilde{\mathcal{O}} \Big(\frac{L}{\mu(K+1)} + \frac{\delta + {}{\mu}}{\mu (1 - \rho)} + \frac{\rho }{(1 - \rho)^2} \cdot \frac{L+ \delta}{\mu}\Big)}, %\zhize{seems to be O~\tilde{\mathcal{O}}} {where KK is the number of additional local update}, ρ\rho measures the network connectivity and δ\delta measures the second-order heterogeneity of the local losses. Our results reveal the tradeoff between communication and computation and show increasing KK can effectively reduce communication costs when the data heterogeneity is low and the network is well-connected. We then consider the over-parameterization regime where the local losses share the same minimums. We proved that employing local updates in DGD, even without gradient correction, achieves exact linear convergence under the Polyak-{\L}ojasiewicz (PL) condition, which can yield a similar effect as DGT in reducing communication complexity. {}{Customization of the result to linear models is further provided, with improved rate expression. }Numerical experiments validate our theoretical results.

Keywords

Cite

@article{arxiv.2403.15654,
  title  = {The Effectiveness of Local Updates for Decentralized Learning under Data Heterogeneity},
  author = {Tongle Wu and Zhize Li and Ying Sun},
  journal= {arXiv preprint arXiv:2403.15654},
  year   = {2024}
}
R2 v1 2026-06-28T15:30:44.200Z