English

Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation

Optimization and Control 2025-01-27 v2 Machine Learning

Abstract

We prove new convergence rates for a generalized version of stochastic Nesterov acceleration under interpolation conditions. Unlike previous analyses, our approach accelerates any stochastic gradient method which makes sufficient progress in expectation. The proof, which proceeds using the estimating sequences framework, applies to both convex and strongly convex functions and is easily specialized to accelerated SGD under the strong growth condition. In this special case, our analysis reduces the dependence on the strong growth constant from ρ\rho to ρ\sqrt{\rho} as compared to prior work. This improvement is comparable to a square-root of the condition number in the worst case and address criticism that guarantees for stochastic acceleration could be worse than those for SGD.

Keywords

Cite

@article{arxiv.2404.02378,
  title  = {Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation},
  author = {Aaron Mishkin and Mert Pilanci and Mark Schmidt},
  journal= {arXiv preprint arXiv:2404.02378},
  year   = {2025}
}

Comments

Warning: this preprint has a significant theoretical bug. We have updated the text to point out the issue and clarify which results are valid

R2 v1 2026-06-28T15:42:29.810Z