Second-Order Information in Non-Convex Stochastic Optimization: Power and Limitations
Abstract
We design an algorithm which finds an -approximate stationary point (with ) using stochastic gradient and Hessian-vector products, matching guarantees that were previously available only under a stronger assumption of access to multiple queries with the same random seed. We prove a lower bound which establishes that this rate is optimal and---surprisingly---that it cannot be improved using stochastic th order methods for any , even when the first derivatives of the objective are Lipschitz. Together, these results characterize the complexity of non-convex stochastic optimization with second-order methods and beyond. Expanding our scope to the oracle complexity of finding -approximate second-order stationary points, we establish nearly matching upper and lower bounds for stochastic second-order methods. Our lower bounds here are novel even in the noiseless case.
Cite
@article{arxiv.2006.13476,
title = {Second-Order Information in Non-Convex Stochastic Optimization: Power and Limitations},
author = {Yossi Arjevani and Yair Carmon and John C. Duchi and Dylan J. Foster and Ayush Sekhari and Karthik Sridharan},
journal= {arXiv preprint arXiv:2006.13476},
year = {2020}
}
Comments
Accepted to CONFERENCE ON LEARNING THEORY (COLT) 2020