English

Trading-off variance and complexity in stochastic gradient descent

Machine Learning 2016-03-23 v1 Information Theory Machine Learning math.IT Optimization and Control

Abstract

Stochastic gradient descent is the method of choice for large-scale machine learning problems, by virtue of its light complexity per iteration. However, it lags behind its non-stochastic counterparts with respect to the convergence rate, due to high variance introduced by the stochastic updates. The popular Stochastic Variance-Reduced Gradient (SVRG) method mitigates this shortcoming, introducing a new update rule which requires infrequent passes over the entire input dataset to compute the full-gradient. In this work, we propose CheapSVRG, a stochastic variance-reduction optimization scheme. Our algorithm is similar to SVRG but instead of the full gradient, it uses a surrogate which can be efficiently computed on a small subset of the input data. It achieves a linear convergence rate ---up to some error level, depending on the nature of the optimization problem---and features a trade-off between the computational complexity and the convergence rate. Empirical evaluation shows that CheapSVRG performs at least competitively compared to the state of the art.

Keywords

Cite

@article{arxiv.1603.06861,
  title  = {Trading-off variance and complexity in stochastic gradient descent},
  author = {Vatsal Shah and Megasthenis Asteris and Anastasios Kyrillidis and Sujay Sanghavi},
  journal= {arXiv preprint arXiv:1603.06861},
  year   = {2016}
}

Comments

14 pages, 13 figures, first edition on 9th of October 2015

R2 v1 2026-06-22T13:16:17.368Z