English

Tight Lower Bounds and Optimal Algorithms for Stochastic Nonconvex Optimization with Heavy-Tailed Noise

Optimization and Control 2026-04-01 v2

Abstract

We study stochastic nonconvex optimization under heavy-tailed noise. In this setting, the stochastic gradients only have bounded pp-th central moment (pp-BCM) for some p(1,2]p \in (1,2]. Building on the foundational work of Arjevani et al. (2022) in stochastic optimization, we establish tight sample complexity lower bounds for all first-order methods under \emph{relaxed} mean-squared smoothness (qq-WAS) and δ\delta-similarity ((q,δ)(q, \delta)-S) assumptions, allowing any exponent q[1,2]q \in [1,2] instead of the standard q=2q = 2. These results substantially broaden the scope of existing lower bounds. To complement them, we show that Normalized Stochastic Gradient Descent with Momentum Variance Reduction (NSGD-MVR), a known algorithm, matches these bounds in expectation. Beyond expectation guarantees, we introduce a new algorithm, Double-Clipped NSGD-MVR, which allows the derivation of high-probability convergence rates under weaker assumptions than in previous works. Finally, for second-order methods with stochastic Hessians satisfying bounded qq-th central moment assumptions for some exponent q[1,2]q \in [1, 2] (allowing qpq \neq p), we establish sharper lower bounds than previous works while improving over Sadiev et al. (2025) (where only p=qp = q is considered) and yielding stronger convergence exponents. Together, these results provide a nearly complete complexity characterization of stochastic nonconvex optimization in heavy-tailed regimes.

Keywords

Cite

@article{arxiv.2512.18713,
  title  = {Tight Lower Bounds and Optimal Algorithms for Stochastic Nonconvex Optimization with Heavy-Tailed Noise},
  author = {Adrien Fradin and Abdurakhmon Sadiev and Laurent Condat and Peter Richtárik},
  journal= {arXiv preprint arXiv:2512.18713},
  year   = {2026}
}