English

First-order algorithms converge faster than $O(1/k)$ on convex problems

Optimization and Control 2019-05-15 v4

Abstract

It is well known that both gradient descent and stochastic coordinate descent achieve a global convergence rate of O(1/k)O(1/k) in the objective value, when applied to a scheme for minimizing a Lipschitz-continuously differentiable, unconstrained convex function. In this work, we improve this rate to o(1/k)o(1/k). We extend the result to proximal gradient and proximal coordinate descent on regularized problems to show similar o(1/k)o(1/k) convergence rates. The result is tight in the sense that a rate of O(1/k1+ϵ)O(1/k^{1+\epsilon}) is not generally attainable for any ϵ>0\epsilon>0, for any of these methods.

Keywords

Cite

@article{arxiv.1812.08485,
  title  = {First-order algorithms converge faster than $O(1/k)$ on convex problems},
  author = {Ching-pei Lee and Stephen J. Wright},
  journal= {arXiv preprint arXiv:1812.08485},
  year   = {2019}
}

Comments

In the proceedings of the 36th International Conference on Machine Learning

R2 v1 2026-06-23T06:51:01.204Z