English
Related papers

Related papers: Central limit theorems for stochastic gradient des…

200 papers

Most existing analyses of (stochastic) gradient descent rely on the condition that for $L$-smooth costs, the step size is less than $2/L$. However, many works have observed that in machine learning applications step sizes often do not…

Optimization and Control · Mathematics 2022-06-10 Kwangjun Ahn , Jingzhao Zhang , Suvrit Sra

The Central Limit Theorem states that, in the limit of a large number of terms, an appropriately scaled sum of independent random variables yields another random variable whose probability distribution tends to a stable distribution. The…

Data Analysis, Statistics and Probability · Physics 2024-04-08 Damián H. Zanette , Inés Samengo

Stochastic gradient descent is one of the most successful approaches for solving large-scale problems, especially in machine learning and statistics. At each iteration, it employs an unbiased estimator of the full gradient computed from one…

Numerical Analysis · Mathematics 2018-12-05 Bangti Jin , Xiliang Lu

Recent theoretical works have characterized the dynamics of wide shallow neural networks trained via gradient descent in an asymptotic mean-field limit when the width tends towards infinity. At initialization, the random sampling of the…

Probability · Mathematics 2022-03-29 Zhengdao Chen , Grant M. Rotskoff , Joan Bruna , Eric Vanden-Eijnden

We consider the optimization of a smooth and strongly convex objective using constant step-size stochastic gradient descent (SGD) and study its properties through the prism of Markov chains. We show that, for unbiased gradient estimates…

Machine Learning · Statistics 2025-11-25 Ibrahim Merad , Stéphane Gaïffas

We consider generalized inversions and descents in finite Weyl groups. We establish Coxeter-theoretic properties of indicator random variables of positive roots such as the covariance of two such indicator random variables. We then compute…

Probability · Mathematics 2023-09-29 Kathrin Meier , Christian Stump

We study the problem of finding the global Riemannian center of mass of a set of data points on a Riemannian manifold. Specifically, we investigate the convergence of constant step-size gradient descent algorithms for solving this problem.…

Differential Geometry · Mathematics 2012-01-05 Bijan Afsari , Roberto Tron , René Vidal

We study the central limit theorem in the non-normal domain of attraction to symmetric $\alpha$-stable laws for $0<\alpha\leq2$. We show that for i.i.d. random variables $X_i$, the convergence rate in $L^\infty$ of both the densities and…

Probability · Mathematics 2018-04-24 Christoph Börgers , Claude Greengard

The global clustering coefficient serves as a powerful metric for the structural analysis and comparison of complex networks. Random geometric graphs offer a realistic framework for representing the spatial constraints and geometry often…

Statistics Theory · Mathematics 2026-02-23 Mingao Yuan , Md. Niamul Islam Sium

A variant of consensus based distributed gradient descent (\textbf{DGD}) is studied for finite sums of smooth but possibly non-convex functions. In particular, the local gradient term in the fixed step-size iteration of each agent is…

Optimization and Control · Mathematics 2026-05-27 Lei Qin , Michael Cantoni , Ye Pu

We prove that the norm version of the adaptive stochastic gradient method (AdaGrad-Norm) achieves a linear convergence rate for a subset of either strongly convex functions or non-convex functions that satisfy the Polyak Lojasiewicz (PL)…

Machine Learning · Statistics 2020-06-23 Yuege Xie , Xiaoxia Wu , Rachel Ward

We establish central limit theorems for the Sample Average Approximation (SAA) method in discrete-time, finite-horizon stochastic optimal control. Our analysis is based on an abstract limit theorem for stochastic backward recursions, which…

Optimization and Control · Mathematics 2026-04-21 Johannes Milz , Alexander Shapiro

Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or ``stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large…

Machine Learning · Computer Science 2025-10-21 Lachlan Ewen MacDonald , Hancheng Min , Leandro Palma , Salma Tarmoun , Ziqing Xu , René Vidal

Polyak-Ruppert averaging yields an asymptotically normal estimator with sandwich covariance $H^{-1}SH^{-1}$, the foundation of online inference. When the gradient step is preconditioned by a data-driven matrix $P_t$, we ask how fast $P_t$…

Statistics Theory · Mathematics 2026-04-28 Sunyoung An , Xiaoming Huo

This paper proposes an asymptotic theory for online inference of the stochastic gradient descent (SGD) iterates with dropout regularization in linear regression. Specifically, we establish the geometric-moment contraction (GMC) for constant…

Machine Learning · Statistics 2024-09-12 Jiaqi Li , Johannes Schmidt-Hieber , Wei Biao Wu

We characterize the convergence in distribution to a standard normal law for a sequence of multiple stochastic integrals of a fixed order with variance converging to 1. Some applications are given, in particular to study the limiting…

Probability · Mathematics 2007-05-23 David Nualart , Giovanni Peccati

We propose and analyze a variant of the classic Polyak-Ruppert averaging scheme, broadly used in stochastic gradient methods. Rather than a uniform average of the iterates, we consider a weighted average, with weights decaying in a…

Machine Learning · Computer Science 2018-02-23 Gergely Neu , Lorenzo Rosasco

We investigate three types of averaging principles and the normal deviation for multi-scale stochastic differential equations (in short, SDEs) with polynomial nonlinearity. More specifically, we first demonstrate the strong convergence of…

Dynamical Systems · Mathematics 2023-08-22 Mengyu Cheng , Zhenxin Liu , Michael Röckner

We study the asymptotic behavior for an inhomogeneous multiscale stochastic dynamical system with non-smooth coefficients. Depending on the averaging regime and the homogenization regime, two strong convergences in the averaging principle…

Probability · Mathematics 2021-04-21 Michael Röckner , Longjie Xie

When training neural networks with low-precision computation, rounding errors often cause stagnation or are detrimental to the convergence of the optimizers; in this paper we study the influence of rounding errors on the convergence of the…

Machine Learning · Statistics 2025-01-22 Lu Xia , Michiel E. Hochstenbach , Stefano Massei