Tight Lower Bounds and Optimal Algorithms for Stochastic Nonconvex Optimization with Heavy-Tailed Noise
Abstract
We study stochastic nonconvex optimization under heavy-tailed noise. In this setting, the stochastic gradients only have bounded -th central moment (-BCM) for some . Building on the foundational work of Arjevani et al. (2022) in stochastic optimization, we establish tight sample complexity lower bounds for all first-order methods under \emph{relaxed} mean-squared smoothness (-WAS) and -similarity (-S) assumptions, allowing any exponent instead of the standard . These results substantially broaden the scope of existing lower bounds. To complement them, we show that Normalized Stochastic Gradient Descent with Momentum Variance Reduction (NSGD-MVR), a known algorithm, matches these bounds in expectation. Beyond expectation guarantees, we introduce a new algorithm, Double-Clipped NSGD-MVR, which allows the derivation of high-probability convergence rates under weaker assumptions than in previous works. Finally, for second-order methods with stochastic Hessians satisfying bounded -th central moment assumptions for some exponent (allowing ), we establish sharper lower bounds than previous works while improving over Sadiev et al. (2025) (where only is considered) and yielding stronger convergence exponents. Together, these results provide a nearly complete complexity characterization of stochastic nonconvex optimization in heavy-tailed regimes.
Cite
@article{arxiv.2512.18713,
title = {Tight Lower Bounds and Optimal Algorithms for Stochastic Nonconvex Optimization with Heavy-Tailed Noise},
author = {Adrien Fradin and Abdurakhmon Sadiev and Laurent Condat and Peter Richtárik},
journal= {arXiv preprint arXiv:2512.18713},
year = {2026}
}