English
Related papers

Related papers: Self-normalized Cram\'er-type Moderate Deviation o…

200 papers

This manuscript studies the Gaussian approximation of the coordinate-wise maximum of self-normalized statistics in high-dimensional settings. We derive an explicit Berry-Esseen bound under weak assumptions on the absolute moments. When the…

Probability · Mathematics 2025-01-16 Woonyoung Chang , Kenta Takatsu , Konrad Urban , Arun Kumar Kuchibhotla

We analyze the implicit bias of constant step stochastic subgradient descent (SGD). We consider the setting of binary classification with homogeneous neural networks - a large class of deep neural networks with ReLU-type activation…

Machine Learning · Computer Science 2025-07-18 Sholom Schechtman , Nicolas Schreuder

Stochastic learning dynamics based on Langevin or Levy stochastic differential equations (SDEs) in deep neural networks control the variance of noise by varying the size of the mini-batch or directly those of injecting noise. Since the…

Machine Learning · Computer Science 2023-10-05 JInwuk Seok , Changsik Cho

We propose a novel approach to numerically approximate McKean-Vlasov stochastic differential equations (MV-SDE) using stochastic gradient descent (SGD) while avoiding the use of interacting particle systems (IPS) {and the associated…

Numerical Analysis · Mathematics 2026-01-22 Ankush Agarwal , Andrea Amato , Goncalo dos Reis , Stefano Pagliarani

We study the approximation of the ergodic measure of the following stochastic differential equation (SDE) on $\mathbb{R}^d$: \begin{eqnarray}\label{e:SDEE} d X_t &=& (b_1(X_t)+b_2(X_t)) d t+\sigma(X_t) d W_t, \end{eqnarray} where $W_t$ is a…

Probability · Mathematics 2023-01-24 Xinghu Jin , Wei Wang , Lihu Xu , Tusheng Zhang

We introduce a constructive framework to learn effective Langevin equations from stationary time series. Unlike conventional approaches that require iterative calibration to match target statistics, our construction guarantees the observed…

Chaotic Dynamics · Physics 2026-02-16 Ludovico Theo Giorgini

The convergence of stochastic interacting particle systems in the mean-field limit to solutions of conservative stochastic partial differential equations is established, with optimal rate of convergence. As a second main result, a…

Probability · Mathematics 2022-12-15 Benjamin Gess , Rishabh S. Gvalani , Vitalii Konarovskyi

In this work, we study the asymptotic randomness of an algorithmic estimator of the saddle point of a globally convex-concave and locally strongly-convex strongly-concave objective. Specifically, we show that the averaged iterates of a…

Optimization and Control · Mathematics 2023-11-07 Abhishek Roy , Yi-An Ma

To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and costly to tune. Existing adaptive strategies based on…

We show that differentially private stochastic gradient descent (DP-SGD) can yield poorly calibrated, overconfident deep learning models. This represents a serious issue for safety-critical applications, e.g. in medical diagnosis. We…

In this paper, we examine the problem of sampling from log-concave distributions with (possibly) superlinear gradient growth under kinetic (underdamped) Langevin algorithms. Using a carefully tailored taming scheme, we propose two novel…

Probability · Mathematics 2025-12-10 Iosif Lytras , Panayotis Mertikopoulos

We propose new limiting dynamics for stochastic gradient descent in the small learning rate regime called stochastic modified flows. These SDEs are driven by a cylindrical Brownian motion and improve the so-called stochastic modified…

Probability · Mathematics 2023-02-15 Benjamin Gess , Sebastian Kassing , Vitalii Konarovskyi

Stochastic Gradient Descent (SGD) has become a cornerstone method in modern data science. However, deploying SGD in high-stakes applications necessitates rigorous quantification of its inherent uncertainty. In this work, we establish…

Machine Learning · Computer Science 2025-10-23 Bhavya Agrawalla , Krishnakumar Balasubramanian , Promit Ghosal

Application of the replica exchange (i.e., parallel tempering) technique to Langevin Monte Carlo algorithms, especially stochastic gradient Langevin dynamics (SGLD), has scored great success in non-convex learning problems, but one…

Numerical Analysis · Mathematics 2023-01-06 Guanxun Li , Guang Lin , Zecheng Zhang , Quan Zhou

We adapt Stein's method to obtain Berry--Esseen type error bounds in the multivariate central limit theorem for non-stationary processes generated by time-dependent compositions of uniformly expanding dynamical systems. In a particular case…

Dynamical Systems · Mathematics 2026-03-17 Juho Leppänen

We develop a framework that allows the use of the multi-level Monte Carlo (MLMC) methodology (Giles2015) to calculate expectations with respect to the invariant measure of an ergodic SDE. In that context, we study the (over-damped) Langevin…

Numerical Analysis · Mathematics 2019-08-13 Michael B. Giles , Mateusz B. Majka , Lukasz Szpruch , Sebastian Vollmer , Konstantinos Zygalakis

his study presents a novel technique to estimate the computational complexity of sequential decoding using the Berry-Esseen theorem. Unlike the theoretical bounds determined by the conventional central limit theorem argument, which often…

Information Theory · Computer Science 2007-08-20 Po-Ning Chen , Yunghsiang S. Han , Carlos R. P. Hartmann , Hong-Bin Wu

We study the self-normalized sums of independent random variables from the perspective of the Malliavin calculus. We give the chaotic expansion for them and we prove a Berry-Ess\'een bound with respect to several distances.

Probability · Mathematics 2014-09-05 Solesne Bourguin , Ciprian Tudor

The stochastic gradient noise (SGN) is a significant factor in the success of stochastic gradient descent (SGD). Following the central limit theorem, SGN was initially modeled as Gaussian, and lately, it has been suggested that stochastic…

Machine Learning · Computer Science 2023-03-07 Barak Battash , Ofir Lindenbaum

While momentum-based accelerated variants of stochastic gradient descent (SGD) are widely used when training machine learning models, there is little theoretical understanding on the generalization error of such methods. In this work, we…

Machine Learning · Computer Science 2024-01-17 Ali Ramezani-Kebrya , Kimon Antonakopoulos , Volkan Cevher , Ashish Khisti , Ben Liang