English
Related papers

Related papers: Error estimates between SGD with momentum and unde…

200 papers

The non-asymptotic analysis of Stochastic Gradient Descent (SGD) typically yields bounds that decompose into a bias term and a variance term. In this work, we focus on the bias component and study the extent to which SGD can match the…

Optimization and Control · Mathematics 2026-02-02 Daniel Cortild , Lucas Ketels , Juan Peypouquet , Guillaume Garrigos

Although generative diffusion models (GDMs) are widely used in practice, their theoretical foundations remain limited, especially concerning the impact of different discretization schemes applied to the underlying stochastic differential…

Numerical Analysis · Mathematics 2026-01-27 Emanuel Pfarr , Radu Timofte , Frank Werner

In this paper we analyze the behaviour of the stochastic gradient descent (SGD), a widely used method in supervised learning for optimizing neural network weights via a minimization of non-convex loss functions. Since the pioneering work of…

Machine Learning · Computer Science 2025-05-13 Davide Barbieri , Matteo Bonforte , Peio Ibarrondo

The $L^k$-Wasserstein distance $\mathbb{W}_k (k\ge 1)$ and the probability distance $\mathbb{W}_\psi$ induced by a concave function $\psi$, are estimated between different diffusion processes with singular coefficients. As applications, the…

Probability · Mathematics 2023-11-07 Xing Huang , Panpan Ren , Feng-Yu Wang

Langevin Monte Carlo (LMC) and its stochastic gradient versions are powerful algorithms for sampling from complex high-dimensional distributions. To sample from a distribution with density $\pi(\theta)\propto \exp(-U(\theta)) $, LMC…

Computation · Statistics 2023-09-25 Sifan Liu

Using the recently published GJF-2GJ Langevin thermostat, which can produce time-step-independent statistical measures even for large time steps, we analyze and discuss the causes for abrupt deviations in statistical data as the time step…

Statistical Mechanics · Physics 2020-01-31 Lucas Frese Grønbech Jensen , Niels Grønbech-Jensen

We study the quantitative convergence of drift-diffusion PDEs that arise as Wasserstein gradient flows of linearly convex functions over the space of probability measures on ${\mathbb R}^d$. In this setting, the objective is in general not…

Optimization and Control · Mathematics 2025-07-17 Lénaïc Chizat , Maria Colombo , Xavier Fernández-Real

The mean-field Langevin dynamics (MFLD) is a nonlinear generalization of the Langevin dynamics that incorporates a distribution-dependent drift, and it naturally arises from the optimization of two-layer neural networks via (noisy) gradient…

Machine Learning · Computer Science 2023-06-13 Taiji Suzuki , Denny Wu , Atsushi Nitanda

In this note we explore how standard statistical distances are equivalent for discrete log-concave distributions. Distances include total variation distance, Wasserstein distance, and $f$-divergences.

Probability · Mathematics 2024-09-10 Arnaud Marsiglietti , Puja Pandey

Stochastic gradient Langevin dynamics (SGLD) and stochastic gradient Hamiltonian Monte Carlo (SGHMC) are two popular Markov Chain Monte Carlo (MCMC) algorithms for Bayesian inference that can scale to large datasets, allowing to sample from…

Machine Learning · Statistics 2021-08-30 Mert Gürbüzbalaban , Xuefeng Gao , Yuanhan Hu , Lingjiong Zhu

Algorithmic stability is an important notion that has proven powerful for deriving generalization bounds for practical algorithms. The last decade has witnessed an increasing number of stability bounds for different algorithms applied on…

Machine Learning · Statistics 2023-10-31 Lingjiong Zhu , Mert Gurbuzbalaban , Anant Raj , Umut Simsekli

While low-precision optimization has been widely used to accelerate deep learning, low-precision sampling remains largely unexplored. As a consequence, sampling is simply infeasible in many large-scale scenarios, despite providing…

Machine Learning · Computer Science 2022-06-22 Ruqi Zhang , Andrew Gordon Wilson , Christopher De Sa

Asynchronous stochastic gradient descent (SGD) enables scalable distributed training but suffers from gradient staleness. Existing mitigation strategies, such as delay-adaptive learning rates and staleness-aware filtering, typically…

Machine Learning · Computer Science 2026-05-15 Tehila Dahan , Roie Reshef , Sharon Goldstein , Kfir Y. Levy

Stochastic Gradient Descent with a constant learning rate (constant SGD) simulates a Markov chain with a stationary distribution. With this perspective, we derive several new results. (1) We show that constant SGD can be used as an…

Machine Learning · Statistics 2018-01-23 Stephan Mandt , Matthew D. Hoffman , David M. Blei

In this paper we propose stochastic gradient-free methods and accelerated methods with momentum for solving stochastic optimization problems. All these methods rely on stochastic directions rather than stochastic gradients. We analyze the…

Optimization and Control · Mathematics 2020-01-15 Xiaopeng Luo , Xin Xu

We investigate the test risk of continuous-time stochastic gradient flow dynamics in learning theory. Using a path integral formulation we provide, in the regime of a small learning rate, a general formula for computing the difference…

Machine Learning · Statistics 2025-03-05 Rodrigo Veiga , Anastasia Remizova , Nicolas Macris

Stochastic natural gradient variational inference (NGVI) is a popular and efficient algorithm for Bayesian inference. Despite empirical success, the convergence of this method is still not fully understood. In this work, we define and study…

Methodology · Statistics 2026-04-02 Thomas Guilmeau , Hadrien Hendrikx , Florence Forbes

Stochastic gradient descent is a classic algorithm that has gained great popularity especially in the last decades as the most common approach for training models in machine learning. While the algorithm has been well-studied when…

Machine Learning · Statistics 2025-09-09 Jose Blanchet , Aleksandar Mijatović , Wenhao Yang

Obtaining coarse-grained models that accurately incorporate finite-size effects is an important open challenge in the study of complex, multi-scale systems. We apply Langevin regression, a recently developed method for finding stochastic…

Adaptation and Self-Organizing Systems · Physics 2021-10-12 Jordan Snyder , Jared L. Callaham , Steven L. Brunton , J. Nathan Kutz

We study a distributed consensus-based stochastic gradient descent (SGD) algorithm and show that the rate of convergence involves the spectral properties of two matrices: the standard spectral gap of a weight matrix from the network…

Optimization and Control · Mathematics 2016-09-02 Avleen S. Bijral , Anand D. Sarwate , Nathan Srebro
‹ Prev 1 8 9 10 Next ›