English
Related papers

Related papers: Central limit theorems for stochastic gradient des…

200 papers

We study the central limit theorem for sums of independent tensor powers, $\frac{1}{\sqrt{d}}\sum\limits_{i=1}^d X_i^{\otimes p}$. We focus on the high-dimensional regime where $X_i \in \mathbb{R}^n$ and $n$ may scale with $d$. Our main…

Probability · Mathematics 2020-11-05 Dan Mikulincer

For a stationary sequence that is regularly varying and associated we give conditions which guarantee that partial sums of this sequence, under normalization related to the exponent of regular variation, converge in distribution to a…

Probability · Mathematics 2019-10-29 Adam Jakubowski

In a vast area of probabilistic limit theorems for dynamical systems with chaotic behaviors always only functional form (exponential, power, etc) of the asymptotic laws and of convergence rates were studied. However, for basically all…

Dynamical Systems · Mathematics 2023-06-28 Leonid A. Bunimovich , Yaofeng Su

Under certain general conditions, we prove that the stable central limit theorem holds in the total variation distance and get its optimal convergence rate for all $\alpha \in (0,2)$. Our method is by two measure decompositions, one step…

Probability · Mathematics 2023-12-08 Xiang Li , Lihu Xu , Haoran Yang

We consider a family of multivariate autoregressive stochastic sequences that restart when hit a neighbourhood of the origin, and study their distributional limits when the autoregressive coefficient tends to one, the noise scaling…

Probability · Mathematics 2020-11-20 Sergey Foss , Matthias Schulte

We extend the central limit theorem under the Dedecker-Rio condition to adapted stationary and ergodic sequences of random variables taking values in a class of smooth Banach spaces. This result applies to the case of random variables…

Probability · Mathematics 2024-07-12 Aurélie Bigot

We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using \emph{any} constant learning rate. This result demonstrates that the stochastic…

Machine Learning · Computer Science 2025-02-12 Jincheng Mei , Bo Dai , Alekh Agarwal , Sharan Vaswani , Anant Raj , Csaba Szepesvari , Dale Schuurmans

We prove quenched versions of (i) a large deviations principle (LDP), (ii) a central limit theorem (CLT), and (iii) a local central limit theorem (LCLT) for non-autonomous dynamical systems. A key advance is the extension of the spectral…

Dynamical Systems · Mathematics 2018-02-14 Davor Dragicevic , Gary Froyland , Cecilia Gonzalez-Tokman , Sandro Vaienti

Most prior work on the convergence of gradient descent (GD) for overparameterized neural networks relies on strong assumptions on the step size (infinitesimal), the hidden-layer width (infinite), or the initialization (large, spectral,…

Machine Learning · Computer Science 2025-05-20 Ziqing Xu , Hancheng Min , Salma Tarmoun , Enrique Mallada , Rene Vidal

We study stochastic gradient descent (SGD) for composite optimization problems with $N$ sequential operators subject to perturbations in both the forward and backward passes. Unlike classical analyses that treat gradient noise as additive…

Optimization and Control · Mathematics 2026-02-25 Boao Kong , Hengrui Zhang , Kun Yuan

We study the convergence rate of randomly truncated stochastic algorithms, which consist in the truncation of the standard Robbins-Monro procedure on an increasing sequence of compact sets. Such a truncation is often required in practice to…

Probability · Mathematics 2010-04-08 Jérôme Lelong

We study the convergence rate of randomly truncated stochastic algorithms, which consist in the truncation of the standard Robbins-Monro procedure on an increasing sequence of compact sets. Such a truncation is often required in practice to…

Probability · Mathematics 2010-03-23 Jérôme Lelong

The growing size of available data has attracted increasing interest in solving minimax problems in a decentralized manner for various machine learning tasks. Previous theoretical research has primarily focused on the convergence rate and…

Machine Learning · Computer Science 2023-11-01 Miaoxi Zhu , Li Shen , Bo Du , Dacheng Tao

We consider the problem of minimizing a non-convex function over a smooth manifold $\mathcal{M}$. We propose a novel algorithm, the Orthogonal Directions Constrained Gradient Method (ODCGM) which only requires computing a projection onto a…

Optimization and Control · Mathematics 2023-03-17 Sholom Schechtman , Daniil Tiapkin , Michael Muehlebach , Eric Moulines

We study a fixed step-size noisy distributed gradient descent algorithm for solving optimization problems in which the objective is a finite sum of smooth but possibly non-convex functions. Random perturbations are introduced to the…

Optimization and Control · Mathematics 2023-07-21 Lei Qin , Michael Cantoni , Ye Pu

We develop a high-dimensional scaling limit for Stochastic Gradient Descent with Polyak Momentum (SGD-M) and adaptive step-sizes. This provides a framework to rigourously compare online SGD with some of its popular variants. We show that…

Machine Learning · Statistics 2026-02-19 Aukosh Jagannath , Taj Jones-McCormick , Varnan Sarangian

This paper presents an algorithmic framework for solving unconstrained stochastic optimization problems using only stochastic function evaluations. We employ central finite-difference based gradient estimation methods to approximate the…

Optimization and Control · Mathematics 2025-01-14 Raghu Bollapragada , Cem Karamanli

We study the generalization properties of unregularized gradient methods applied to separable linear classification -- a setting that has received considerable attention since the pioneering work of Soudry et al. (2018). We establish tight…

Machine Learning · Computer Science 2023-03-03 Matan Schliserman , Tomer Koren

Randomized coordinate descent (RCD) is a popular optimization algorithm with wide applications in solving various machine learning problems, which motivates a lot of theoretical analysis on its convergence behavior. As a comparison, there…

Machine Learning · Computer Science 2021-08-18 Puyu Wang , Liang Wu , Yunwen Lei

We investigate statistical properties of the optimal value of the Sample Average Approximation of stochastic programs, continuing the study in Kr\"atschmer (2023). Central Limit Theorem type results are derived for the optimal value. As a…

Optimization and Control · Mathematics 2023-12-12 Volker Krätschmer