English

Convergence Rate for the Last Iterate of Stochastic Gradient Descent Schemes

Optimization and Control 2026-03-11 v4 Machine Learning

Abstract

We study the convergence rate for the last iterate of stochastic gradient descent (SGD) and stochastic heavy ball (SHB) in the parametric setting when the objective function FF is globally convex or non-convex whose gradient is γ\gamma-H\"{o}lder. Using only discrete Gronwall's inequality without Robbins-Siegmund theorem, we recover results for both SGD and SHB: minstF(ws)2=o(tp1)\min_{s\leq t} \|\nabla F(w_s)\|^2 = o(t^{p-1}) for non-convex objectives and F(wτt)F=o(t2γ/(1+γ)max(p1,2p+1)ϵ)F(w_{\tau \wedge t}) - F_* = o(t^{2\gamma/(1+\gamma) \cdot \max(p-1,-2p+1)-\epsilon}) for β(0,1)\beta \in (0, 1), τ:=inf{t>0:F(wt)=F}\tau := \inf \{ t > 0 : F(w_t) = F_*\}, and minstF(ws)F=o(tp1)\min_{s \leq t} F(w_s) - F_* = o(t^{p-1}) for convex objectives FF whose minimum is FF_*. In addition, we proved that SHB with constant momentum parameter β(0,1)\beta \in (0, 1) attains a convergence rate of F(wt)F=O(tmax(p1,2p+1)log2tδ)F(w_t) - F_* = O(t^{\max(p-1,-2p+1)} \log^2 \frac{t}{\delta}) with probability at least 1δ1-\delta when FF is convex and γ=1\gamma = 1 and step size αt=Θ(tp)\alpha_t = \Theta(t^{-p}) with p(12,1)p \in (\frac{1}{2}, 1).

Keywords

Cite

@article{arxiv.2507.07281,
  title  = {Convergence Rate for the Last Iterate of Stochastic Gradient Descent Schemes},
  author = {Marcel Hudiani},
  journal= {arXiv preprint arXiv:2507.07281},
  year   = {2026}
}