Related papers: Uniform-in-Time Weak Propagation-of-Chaos in Shall…
We study the kinetic mean field Langevin dynamics under the functional convexity assumption of the mean field energy functional. Using hypocoercivity, we first establish the exponential convergence of the mean field dynamics and then show…
We study the uniform-in-time weak propagation of chaos for the consensus-based optimization (CBO) method on a bounded searching domain. We apply the methodology for studying long-time behaviors of interacting particle systems developed in…
We study the convergence of gradient methods for the training of mean-field single-hidden-layer neural networks with square loss. For this high-dimensional and non-convex optimization problem, most known convergence results are either…
We propose a custom learning algorithm for shallow over-parameterized neural networks, i.e., networks with single hidden layer having infinite width. The infinite width of the hidden layer serves as an abstraction for the…
We study the asymptotic behavior, uniform-in-time, of a non-linear dynamical system under the combined effects of fast periodic sampling with period $\delta$ and small white noise of size $\varepsilon,\thinspace 0<\varepsilon,\delta \ll 1$.…
In this paper, we investigate the limiting behavior of a continuous-time counterpart of the Stochastic Gradient Descent (SGD) algorithm applied to two-layer overparameterized neural networks, as the number or neurons (ie, the size of the…
We quantify, uniformly over time and with high probability, the discrepancy between the predictions of a two-layer neural network trained by stochastic gradient descent (SGD) and their mean-field limit, for quadratic loss and ridge…
We study the approximation gap between the dynamics of a polynomial-width neural network and its infinite-width counterpart, both trained using projected gradient descent in the mean-field scaling regime. We demonstrate how to tightly bound…
We consider shallow (single hidden layer) neural networks and characterize their performance when trained with stochastic gradient descent as the number of hidden units $N$ and gradient descent steps grow to infinity. In particular, we…
In this paper, uniform in time quantitative propagation of chaos in $L^1$-Wasserstein distance for mean field interacting particle system is derived, where the diffusion coefficient is allowed to be interacting and the drift is assumed to…
The mean field (MF) theory of multilayer neural networks centers around a particular infinite-width scaling, where the learning dynamics is closely tracked by the MF limit. A random fluctuation around this infinite-width limit is expected…
Based on a coupling approach, we prove uniform in time propagation of chaos for weakly interacting mean-field particle systems with possibly non-convex confinement and interaction potentials. The approach is based on a combination of…
We address the long time behaviour of weakly interacting diffusive particle systems on the d-dimensional torus. Our main result is to show that, under certain regularity conditions, the weak error between the empirical distribution of the…
We rigorously prove a central limit theorem for neural network models with a single hidden layer. The central limit theorem is proven in the asymptotic regime of simultaneously (A) large numbers of hidden units and (B) large numbers of…
Mean-field Langevin dynamics (MFLD) minimizes an entropy-regularized nonlinear convex functional defined over the space of probability distributions. MFLD has gained attention due to its connection with noisy gradient descent for mean-field…
Time-uniform log-Sobolev inequalities (LSI) satisfied by solutions of semi-linear mean-field equations have recently appeared to be a key tool to obtain time-uniform propagation of chaos estimates. This work addresses the more general…
In this article, we are interested in the behavior of a fully connected network of $N$ neurons, where $N$ tends to infinity. We assume that the neurons follow the stochastic FitzHugh-Nagumo model, whose specificity is the non-linearity with…
In this paper, we provide the first precise distributional characterization of gradient descent iterates for general multi-layer neural networks under the canonical single-index regression model, in the `finite-width proportional regime'…
Algorithmic stability is an important notion that has proven powerful for deriving generalization bounds for practical algorithms. The last decade has witnessed an increasing number of stability bounds for different algorithms applied on…
We study norm-based uniform convergence bounds for neural networks, aiming at a tight understanding of how these are affected by the architecture and type of norm constraint, for the simple class of scalar-valued one-hidden-layer networks,…