English

High-probability Bounds for Non-Convex Stochastic Optimization with Heavy Tails

Machine Learning 2021-11-10 v2 Optimization and Control Machine Learning

Abstract

We consider non-convex stochastic optimization using first-order algorithms for which the gradient estimates may have heavy tails. We show that a combination of gradient clipping, momentum, and normalized gradient descent yields convergence to critical points in high-probability with best-known rates for smooth losses when the gradients only have bounded p\mathfrak{p}th moments for some p(1,2]\mathfrak{p}\in(1,2]. We then consider the case of second-order smooth losses, which to our knowledge have not been studied in this setting, and again obtain high-probability bounds for any p\mathfrak{p}. Moreover, our results hold for arbitrary smooth norms, in contrast to the typical SGD analysis which requires a Hilbert space norm. Further, we show that after a suitable "burn-in" period, the objective value will monotonically decrease for every iteration until a critical point is identified, which provides intuition behind the popular practice of learning rate "warm-up" and also yields a last-iterate guarantee.

Keywords

Cite

@article{arxiv.2106.14343,
  title  = {High-probability Bounds for Non-Convex Stochastic Optimization with Heavy Tails},
  author = {Ashok Cutkosky and Harsh Mehta},
  journal= {arXiv preprint arXiv:2106.14343},
  year   = {2021}
}
R2 v1 2026-06-24T03:38:53.295Z