Related papers: Central limit theorems for stochastic gradient des…
We show that gradient descent converges to a local minimizer, almost surely with random initialization. This is proved by applying the Stable Manifold Theorem from dynamical systems theory.
Optimization algorithms are unlikely to converge to strict saddle points. Proofs to that effect rely on the Center-Stable Manifold Theorem (CSMT), casting algorithms as dynamical systems: $x_{k+1} = g_k(x_k)$. In its standard form, the CSMT…
We prove quenched versions of a central limit theorem, a large deviations principle as well as a local central limit theorem for expanding on average cocycles. This is achieved by building an appropriate modification of the spectral method…
In this paper, we consider a generalization of the elephant random walk model. Compared to the usual elephant random walk, an interesting feature of this model is that the step sizes form a sequence of positive independent and identically…
This paper presents a novel stochastic gradient descent algorithm for constrained optimization. The proposed algorithm randomly samples constraints and components of the finite sum objective function and relies on a relaxed logarithmic…
The main result of this paper is a general central limit theorem for distributions defined by certain renewal type equations. We apply this to weakly self-avoiding random walks. We give good error estimates and Gaussian tail estimates which…
In this paper, we establish a Quantitative Central Limit Theorem ({\sc qclt}) for the Stochastic Gradient Descent in Continuous Time ({\sc sgdct}) algorithm, whose parameter updates are governed by a stochastic differential equation. We…
This paper investigates asymptotic behaviors of gradient descent algorithms (particularly accelerated gradient descent and stochastic gradient descent) in the context of stochastic optimization arising in statistics and machine learning…
The purpose of this work is to establish a central limit theorem that can be applied to a particular form of Markov chains, including the number of descents in a random permutation of $\mathfrak{S}_n$, two-type generalized P{\'o}lya urns,…
The paper studies a distributed gradient descent (DGD) process and considers the problem of showing that in nonconvex optimization problems, DGD typically converges to local minima rather than saddle points. The paper considers…
In this work, we study the asymptotic randomness of an algorithmic estimator of the saddle point of a globally convex-concave and locally strongly-convex strongly-concave objective. Specifically, we show that the averaged iterates of a…
Stochastic optimization naturally appear in many application areas, including machine learning. Our goal is to go further in the analysis of the Stochastic Average Gradient Accelerated (SAGA) algorithm. To achieve this, we introduce a new…
We rigorously prove a central limit theorem for neural network models with a single hidden layer. The central limit theorem is proven in the asymptotic regime of simultaneously (A) large numbers of hidden units and (B) large numbers of…
We show central limit theorems (CLT) for the Stieltjes transforms or more general analytic functions of symmetric matrices with independent heavy tailed entries, including entries in the domain of attraction of $\alpha$-stable laws and…
In this paper, we study the convergence for solutions to a sequence of (possibly degenerate) stochastic differential equations with jumps, when the coefficients converge in some appropriate sense. Our main tools are the superposition…
The convergence rate of stochastic gradient search is analyzed in this paper. Using arguments based on differential geometry and Lojasiewicz inequalities, tight bounds on the convergence rate of general stochastic gradient algorithms are…
We give a new proof of the classical Central Limit Theorem, in the Mallows ($L^r$-Wasserstein) distance. Our proof is elementary in the sense that it does not require complex analysis, but rather makes use of a simple subadditive inequality…
In this article, we consider convergence of stochastic gradient descent schemes (SGD), including momentum stochastic gradient descent (MSGD), under weak assumptions on the underlying landscape. More explicitly, we show that on the event…
Adam is a popular variant of stochastic gradient descent for finding a local minimizer of a function. In the constant stepsize regime, assuming that the objective function is differentiable and non-convex, we establish the convergence in…
Motivated by the stochastic block model, we investigate a class of Wigner-type matrices with certain block structures, and establish a CLT for the corresponding linear spectral statistics via the large-deviation bounds from local law and…