English
Related papers

Related papers: When Does Dynamic Preconditioning Preserve the Pol…

200 papers

Exponential generalization bounds with near-tight rates have recently been established for uniformly stable learning algorithms. The notion of uniform stability, however, is stringent in the sense that it is invariant to the data-generating…

Machine Learning · Statistics 2022-06-09 Xiao-Tong Yuan , Ping Li

This paper proposes an asymptotic theory for online inference of the stochastic gradient descent (SGD) iterates with dropout regularization in linear regression. Specifically, we establish the geometric-moment contraction (GMC) for constant…

Machine Learning · Statistics 2024-09-12 Jiaqi Li , Johannes Schmidt-Hieber , Wei Biao Wu

We study local linear convergence of gradient descent for finite-width feedforward networks under the squared empirical loss. Prior work shows that GD can remain confined to a Locally Quasi-Convex Region (LQCR) around initialization, but…

Machine Learning · Statistics 2026-05-29 Agnideep Aich , Ashit Baran Aich , Bruce Wade

An influential line of recent work has focused on the generalization properties of unregularized gradient-based learning procedures applied to separable linear classification with exponentially-tailed loss functions. The ability of such…

Machine Learning · Computer Science 2022-06-24 Matan Schliserman , Tomer Koren

In this paper, we refine the Berry-Esseen bounds for the multivariate normal approximation of Polyak-Ruppert averaged iterates arising from the linear stochastic approximation (LSA) algorithm with decreasing step size. We consider the…

Machine Learning · Statistics 2025-10-15 Bogdan Butyrin , Eric Moulines , Alexey Naumov , Sergey Samsonov , Qi-Man Shao , Zhuo-Song Zhang

Stochastic gradient descent (SGD) with mini-batching is a standard tool in large-scale optimization, yet its theoretical properties under heavy-tailed gradient noise remain largely unexplored. In this paper we study SGD with increasing…

Probability · Mathematics 2026-05-11 Bartosz Glowacki , Rafal Kulik , Philippe Soulier

To describe the slow dynamics of a system out of equilibrium, but close to a dynamical arrest, we generalize the ideas of previous work to the case where time-translational invariance is broken. We introduce a model of the dynamics that is…

Disordered Systems and Neural Networks · Physics 2016-08-31 P. De Gregorio , F. Sciortino , P. Tartaglia , E. Zaccarelli , K. A. Dawson

Many machine learning and optimization algorithms are built upon the framework of stochastic approximation (SA), for which the selection of step-size (or learning rate) $\{\alpha_n\}$ is crucial for success. An essential condition for…

Statistics Theory · Mathematics 2025-08-05 Caio Kalil Lauand , Sean Meyn

We present bounded dynamic (but observer-free) output feedback laws that achieve global stabilization of equilibrium profiles of the partial differential equation (PDE) model of a simplified, age-structured chemostat model. The chemostat…

Optimization and Control · Mathematics 2016-09-30 Iasson Karafyllis , Miroslav Krstic

We study the asymptotic shape of the trajectory of the stochastic gradient descent algorithm applied to a convex objective function. Under mild regularity assumptions, we prove a functional central limit theorem for the properly rescaled…

Machine Learning · Statistics 2026-02-18 Kessang Flamand , Victor-Emmanuel Brunel

In this work, we show that uniform integrability is not a necessary condition for central limit theorems (CLT) to hold for normalized multilevel Monte Carlo (MLMC) estimators and we provide near optimal weaker conditions under which the CLT…

Probability · Mathematics 2019-05-17 Håkon Hoel , Sebastian Krumscheid

This paper is devoted to the stability analysis of a classical three-field formulation of Biot's consolidation model where the unknown variables are the displacements, fluid flux (Darcy velocity), and pore pressure. Specific…

Numerical Analysis · Mathematics 2018-06-21 Qingguo Hong , Johannes Kraus

Most prior work on the convergence of gradient descent (GD) for overparameterized neural networks relies on strong assumptions on the step size (infinitesimal), the hidden-layer width (infinite), or the initialization (large, spectral,…

Machine Learning · Computer Science 2025-05-20 Ziqing Xu , Hancheng Min , Salma Tarmoun , Enrique Mallada , Rene Vidal

A generalized Cahn-Hilliard model in a bounded interval of the real line with no-flux boundary conditions is considered. The label "generalized" refers to the fact that we consider a concentration dependent mobility, the $p$-Laplace…

Analysis of PDEs · Mathematics 2024-05-20 Raffaele Folino , Luis Fernando Lopez Rios , Marta Strani

Stochastic gradient descent in continuous time (SGDCT) provides a computationally efficient method for the statistical learning of continuous-time models, which are widely used in science, engineering, and finance. The SGDCT algorithm…

Probability · Mathematics 2019-06-18 Justin Sirignano , Konstantinos Spiliopoulos

We consider non-convex stochastic optimization using first-order algorithms for which the gradient estimates may have heavy tails. We show that a combination of gradient clipping, momentum, and normalized gradient descent yields convergence…

Machine Learning · Computer Science 2021-11-10 Ashok Cutkosky , Harsh Mehta

We investigate a specific reaction-diffusion system that admits a monostable pulled front propagating at constant critical speed. When a small parameter changes sign, the stable equilibrium behind the front destabilizes, due to essential…

Analysis of PDEs · Mathematics 2021-10-07 Louis Garénaux

Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stability, but its placement has often been ad-hoc. In this paper, we conduct a principled study on the…

Machine Learning · Computer Science 2025-10-14 Kelvin Kan , Xingjian Li , Benjamin J. Zhang , Tuhin Sahai , Stanley Osher , Krishna Kumar , Markos A. Katsoulakis

In this paper, we establish the central limit theorem (CLT) for the linear spectral statistics (LSS) of sample correlation matrix $R$, constructed from a $p\times n$ data matrix $X$ with independent and identically distributed (i.i.d.)…

Probability · Mathematics 2024-09-20 Yanpeng Li , Guangming Pan , Jiahui Xie , Wang Zhou

High-dimensional Kronecker-structured estimation faces a conflict between non-convex scaling ambiguities and statistical robustness. The arbitrary factor scaling distorts gradient magnitudes, rendering standard fixed-threshold robust…

Methodology · Statistics 2025-12-23 Xiaoyu Zhang , Zhiyun Fan , Wenyang Zhang , Di Wang