Related papers: Central limit theorems for stochastic gradient des…
Stochastic gradient descent is a classic algorithm that has gained great popularity especially in the last decades as the most common approach for training models in machine learning. While the algorithm has been well-studied when…
This paper deals with the numerical approximation of normalizing constants produced by particle methods, in the general framework of Feynman-Kac sequences of measures. It is well-known that the corresponding estimates satisfy a central…
The main objective of this article is to establish a central limit theorem for additive three-variable functionals of bifurcating Markov chains. We thus extend the central limit theorem under point-wise ergodic conditions studied in…
The non-asymptotic analysis of Stochastic Gradient Descent (SGD) typically yields bounds that decompose into a bias term and a variance term. In this work, we focus on the bias component and study the extent to which SGD can match the…
Linear structural error-in-variables models with univariate observations are revisited for studying modified least squares estimators of the slope and intercept. New marginal central limit theorems (CLT's) are established for these…
We give a general local central limit theorem for the sum of two independent random variables, one of which satisfies a central limit theorem while the other satisfies a local central limit theorem with the same order variance. We apply…
We establish central limit theorems for a large class of supercritical branching Markov processes in infinite dimension with spatially dependent and non-necessarily local branching mechanisms. This result relies on a fourth moment…
We prove a central limit theorem applicable to one dimensional stochastic approximation algorithms that converge to a point where the error terms of the algorithm do not vanish. We show how this applies to a certain class of these…
In this article, we quantify the functional convergence of the rescaled random walk with heavy tails to a stable process.This generalizes the Generalized Central Limit Theorem for stable random variables infinite dimension. We show that…
Accelerated gradient descent iterations are widely used in optimization. It is known that, in the continuous-time limit, these iterations converge to a second-order differential equation which we refer to as the accelerated gradient flow.…
For a L\'evy basis $L$ on $\mathbb{R}^d$ and a suitable kernel function $f:\mathbb{R}^d \to \mathbb{R}$, consider the continuous spatial moving average field $X=(X_t)_{t\in \mathbb{R}^d}$ defined by $X_t = \int_{\mathbb{R}^d} f(t-s) \,…
This work focuses on the temporal average of the backward Euler--Maruyama (BEM) method, which is used to approximate the ergodic limit of stochastic ordinary differential equations with super-linearly growing drift coefficients. We give the…
The convergence of stochastic gradient descent is highly dependent on the step-size, especially on non-convex problems such as neural network training. Step decay step-size schedules (constant and then cut) are widely used in practice…
We establish a central limit theorem and prove a moderate deviation principle for stochastic scalar conservation laws. Due to the lack of viscous term, this is done in the framework of kinetic solution. The weak convergence method and…
Gradient clipping is a popular modification to standard (stochastic) gradient descent, at every iteration limiting the gradient norm to a certain value $c >0$. It is widely used for example for stabilizing the training of deep learning…
This paper analyzes the trajectories of stochastic gradient descent (SGD) to help understand the algorithm's convergence properties in non-convex problems. We first show that the sequence of iterates generated by SGD remains bounded and…
This paper presents the asymptotic theory for nondegenerate $U$-statistics of high frequency observations of continuous It\^{o} semimartingales. We prove uniform convergence in probability and show a functional stable central limit theorem…
The stability and generalization of stochastic gradient-based methods provide valuable insights into understanding the algorithmic performance of machine learning models. As the main workhorse for deep learning, stochastic gradient descent…
In this paper, employing the weak convergence method, based on a variational representation for expected values of positive functionals of a Brownian motion, we investigate moderate deviation %(CLT for abbreviation) for a class of…
This paper provides central limit theorems for the wavelet packet decomposition of stationary band-limited random processes. The asymptotic analysis is performed for the sequences of the wavelet packet coefficients returned at the nodes of…