A Newton-Based Method for Nonconvex Optimization with Fast Evasion of Saddle Points
Abstract
Machine learning problems such as neural network training, tensor decomposition, and matrix factorization, require local minimization of a nonconvex function. This local minimization is challenged by the presence of saddle points, of which there can be many and from which descent methods may take inordinately large number of iterations to escape. This paper presents a second-order method that modifies the update of Newton's method by replacing the negative eigenvalues of the Hessian by their absolute values and uses a truncated version of the resulting matrix to account for the objective's curvature. The method is shown to escape saddles in at most iterations where is the target optimality and characterizes a point sufficiently far away from the saddle. This base of this exponential escape is independently of problem constants. Adding classical properties of Newton's method, the paper proves convergence to a local minimum with probability in iterations.
Cite
@article{arxiv.1707.08028,
title = {A Newton-Based Method for Nonconvex Optimization with Fast Evasion of Saddle Points},
author = {Santiago Paternain and Aryan Mokhtari and Alejandro Ribeiro},
journal= {arXiv preprint arXiv:1707.08028},
year = {2018}
}