English

Escaping Saddles with Stochastic Gradients

Machine Learning 2018-09-18 v2 Optimization and Control Machine Learning

Abstract

We analyze the variance of stochastic gradients along negative curvature directions in certain non-convex machine learning models and show that stochastic gradients exhibit a strong component along these directions. Furthermore, we show that - contrary to the case of isotropic noise - this variance is proportional to the magnitude of the corresponding eigenvalues and not decreasing in the dimensionality. Based upon this observation we propose a new assumption under which we show that the injection of explicit, isotropic noise usually applied to make gradient descent escape saddle points can successfully be replaced by a simple SGD step. Additionally - and under the same condition - we derive the first convergence rate for plain SGD to a second-order stationary point in a number of iterations that is independent of the problem dimension.

Keywords

Cite

@article{arxiv.1803.05999,
  title  = {Escaping Saddles with Stochastic Gradients},
  author = {Hadi Daneshmand and Jonas Kohler and Aurelien Lucchi and Thomas Hofmann},
  journal= {arXiv preprint arXiv:1803.05999},
  year   = {2018}
}
R2 v1 2026-06-23T00:54:53.103Z