Related papers: Central limit theorems for stochastic gradient des…
We provide a framework to analyze the convergence of discretized kinetic Langevin dynamics for $M$-$\nabla$Lipschitz, $m$-convex potentials. Our approach gives convergence rates of $\mathcal{O}(m/M)$, with explicit stepsize restrictions,…
In this work, we obtain the central limit theorem for fluctuations of Young diagrams around their limit shape in the bulk of the "spectrum" of partitions of a large integer n (under the Plancherel measure). More specifically, we show that,…
We establish convergence theorems for Riemannian stochastic gradient descents in which the underlying probability spaces vary from iteration to iteration. As applications, we deduce convergence results for Riemannian stochastic gradient…
In this article we establish two fundamental results for the sublevel set persistent homology for stationary processes indexed by the positive integers. The first is a strong law of large numbers for the persistence diagram (treated as a…
We show how a central limit theorem for Poisson model random polygons implies a central limit theorem for uniform model random polygons. To prove this implication, it suffices to show that in the two models, the variables in question have…
Stochastic Gradient Descent (SGD) and its Ruppert-Polyak averaged variant (ASGD) lie at the heart of modern large-scale learning, yet their theoretical properties in high-dimensional settings are rarely understood. In this paper, we provide…
Algorithmic stability is among the most potent techniques in generalization analysis. However, its derivation usually requires a stepsize $\eta_t = \mathcal{O}(1/t)$ under non-convex training regimes, where $t$ denotes iterations. This…
We establish a scaling limit for autonomous stochastic Newton equations, the solutions are often called nonlinear stochastic oscillators, where the nonlinear drift includes a mean field term of McKean type and the driving noise is Gaussian.…
The distribution of descents in fixed conjugacy classes of $S_n$ has been studied, and it is shown that its moments have interesting properties. Kim and Lee showed, by using Curtiss' theorem and moment generating functions, how to prove a…
Stochastic gradient descent (SGD) is a foundational algorithm for large-scale statistical learning and stochastic optimization. However, statistical inference based on SGD iterates remains challenging when stochastic gradients have infinite…
We refine the classical Lindeberg-Feller central limit theorem by obtaining asymptotic bounds on the Kolmogorov distance, the Wasserstein distance, and the parametrized Prokhorov distances in terms of a Lindeberg index. We thus obtain more…
This paper presents some limit theorems for certain functionals of moving averages of semimartingales plus noise which are observed at high frequency. Our method generalizes the pre-averaging approach (see [Bernoulli 15 (2009) 634--658,…
In this paper, we prove a central limit theorem and estabilish a moderate deviation principle for stochastic models of incompressible second fluids. The weak convergence method inreoduced by [4] plays an important role.
The stratified resampling mechanism is one of the resampling schemes commonly used in the resampling steps of particle filters. In the present paper, we prove a central limit theorem for this mechanism under the assumption that the initial…
In these notes, we obtain new stability estimates for centered non-degenerate selfdecomposable probability measures on $\mathbb{R}^d$ with finite second moment and for non-degenerate symmetric $\alpha$-stable probability measures on…
We prove explicit bounds on the exponential rate of convergence for the momentum stochastic gradient descent scheme (MSGD) for arbitrary, fixed hyperparameters (learning rate, friction parameter) and its continuous-in-time counterpart in…
In this paper, we study the convergence properties of the Stochastic Gradient Descent (SGD) method for finding a stationary point of a given objective function $J(\cdot)$. The objective function is not required to be convex. Rather, our…
This paper proposes a novel proximal-gradient algorithm for a decentralized optimization problem with a composite objective containing smooth and non-smooth terms. Specifically, the smooth and nonsmooth terms are dealt with by gradient and…
The main aim of this paper is to provide an analysis of gradient descent (GD) algorithms with gradient errors that do not necessarily vanish, asymptotically. In particular, sufficient conditions are presented for both stability (almost sure…
Recent studies have provided both empirical and theoretical evidence illustrating that heavy tails can emerge in stochastic gradient descent (SGD) in various scenarios. Such heavy tails potentially result in iterates with diverging…