Related papers: Central limit theorems for stochastic gradient des…
We study the central limit theorem for sums of independent tensor powers, $\frac{1}{\sqrt{d}}\sum\limits_{i=1}^d X_i^{\otimes p}$. We focus on the high-dimensional regime where $X_i \in \mathbb{R}^n$ and $n$ may scale with $d$. Our main…
For a stationary sequence that is regularly varying and associated we give conditions which guarantee that partial sums of this sequence, under normalization related to the exponent of regular variation, converge in distribution to a…
In a vast area of probabilistic limit theorems for dynamical systems with chaotic behaviors always only functional form (exponential, power, etc) of the asymptotic laws and of convergence rates were studied. However, for basically all…
Under certain general conditions, we prove that the stable central limit theorem holds in the total variation distance and get its optimal convergence rate for all $\alpha \in (0,2)$. Our method is by two measure decompositions, one step…
We consider a family of multivariate autoregressive stochastic sequences that restart when hit a neighbourhood of the origin, and study their distributional limits when the autoregressive coefficient tends to one, the noise scaling…
We extend the central limit theorem under the Dedecker-Rio condition to adapted stationary and ergodic sequences of random variables taking values in a class of smooth Banach spaces. This result applies to the case of random variables…
We provide a new understanding of the stochastic gradient bandit algorithm by showing that it converges to a globally optimal policy almost surely using \emph{any} constant learning rate. This result demonstrates that the stochastic…
We prove quenched versions of (i) a large deviations principle (LDP), (ii) a central limit theorem (CLT), and (iii) a local central limit theorem (LCLT) for non-autonomous dynamical systems. A key advance is the extension of the spectral…
Most prior work on the convergence of gradient descent (GD) for overparameterized neural networks relies on strong assumptions on the step size (infinitesimal), the hidden-layer width (infinite), or the initialization (large, spectral,…
We study stochastic gradient descent (SGD) for composite optimization problems with $N$ sequential operators subject to perturbations in both the forward and backward passes. Unlike classical analyses that treat gradient noise as additive…
We study the convergence rate of randomly truncated stochastic algorithms, which consist in the truncation of the standard Robbins-Monro procedure on an increasing sequence of compact sets. Such a truncation is often required in practice to…
We study the convergence rate of randomly truncated stochastic algorithms, which consist in the truncation of the standard Robbins-Monro procedure on an increasing sequence of compact sets. Such a truncation is often required in practice to…
The growing size of available data has attracted increasing interest in solving minimax problems in a decentralized manner for various machine learning tasks. Previous theoretical research has primarily focused on the convergence rate and…
We consider the problem of minimizing a non-convex function over a smooth manifold $\mathcal{M}$. We propose a novel algorithm, the Orthogonal Directions Constrained Gradient Method (ODCGM) which only requires computing a projection onto a…
We study a fixed step-size noisy distributed gradient descent algorithm for solving optimization problems in which the objective is a finite sum of smooth but possibly non-convex functions. Random perturbations are introduced to the…
We develop a high-dimensional scaling limit for Stochastic Gradient Descent with Polyak Momentum (SGD-M) and adaptive step-sizes. This provides a framework to rigourously compare online SGD with some of its popular variants. We show that…
This paper presents an algorithmic framework for solving unconstrained stochastic optimization problems using only stochastic function evaluations. We employ central finite-difference based gradient estimation methods to approximate the…
We study the generalization properties of unregularized gradient methods applied to separable linear classification -- a setting that has received considerable attention since the pioneering work of Soudry et al. (2018). We establish tight…
Randomized coordinate descent (RCD) is a popular optimization algorithm with wide applications in solving various machine learning problems, which motivates a lot of theoretical analysis on its convergence behavior. As a comparison, there…
We investigate statistical properties of the optimal value of the Sample Average Approximation of stochastic programs, continuing the study in Kr\"atschmer (2023). Central Limit Theorem type results are derived for the optimal value. As a…