中文
相关论文

相关论文: Error estimates between SGD with momentum and unde…

200 篇论文

The non-asymptotic analysis of Stochastic Gradient Descent (SGD) typically yields bounds that decompose into a bias term and a variance term. In this work, we focus on the bias component and study the extent to which SGD can match the…

最优化与控制 · 数学 2026-02-02 Daniel Cortild , Lucas Ketels , Juan Peypouquet , Guillaume Garrigos

Although generative diffusion models (GDMs) are widely used in practice, their theoretical foundations remain limited, especially concerning the impact of different discretization schemes applied to the underlying stochastic differential…

数值分析 · 数学 2026-01-27 Emanuel Pfarr , Radu Timofte , Frank Werner

In this paper we analyze the behaviour of the stochastic gradient descent (SGD), a widely used method in supervised learning for optimizing neural network weights via a minimization of non-convex loss functions. Since the pioneering work of…

机器学习 · 计算机科学 2025-05-13 Davide Barbieri , Matteo Bonforte , Peio Ibarrondo

The $L^k$-Wasserstein distance $\mathbb{W}_k (k\ge 1)$ and the probability distance $\mathbb{W}_\psi$ induced by a concave function $\psi$, are estimated between different diffusion processes with singular coefficients. As applications, the…

概率论 · 数学 2023-11-07 Xing Huang , Panpan Ren , Feng-Yu Wang

Langevin Monte Carlo (LMC) and its stochastic gradient versions are powerful algorithms for sampling from complex high-dimensional distributions. To sample from a distribution with density $\pi(\theta)\propto \exp(-U(\theta)) $, LMC…

统计计算 · 统计学 2023-09-25 Sifan Liu

Using the recently published GJF-2GJ Langevin thermostat, which can produce time-step-independent statistical measures even for large time steps, we analyze and discuss the causes for abrupt deviations in statistical data as the time step…

统计力学 · 物理学 2020-01-31 Lucas Frese Grønbech Jensen , Niels Grønbech-Jensen

We study the quantitative convergence of drift-diffusion PDEs that arise as Wasserstein gradient flows of linearly convex functions over the space of probability measures on ${\mathbb R}^d$. In this setting, the objective is in general not…

最优化与控制 · 数学 2025-07-17 Lénaïc Chizat , Maria Colombo , Xavier Fernández-Real

The mean-field Langevin dynamics (MFLD) is a nonlinear generalization of the Langevin dynamics that incorporates a distribution-dependent drift, and it naturally arises from the optimization of two-layer neural networks via (noisy) gradient…

机器学习 · 计算机科学 2023-06-13 Taiji Suzuki , Denny Wu , Atsushi Nitanda

In this note we explore how standard statistical distances are equivalent for discrete log-concave distributions. Distances include total variation distance, Wasserstein distance, and $f$-divergences.

概率论 · 数学 2024-09-10 Arnaud Marsiglietti , Puja Pandey

Stochastic gradient Langevin dynamics (SGLD) and stochastic gradient Hamiltonian Monte Carlo (SGHMC) are two popular Markov Chain Monte Carlo (MCMC) algorithms for Bayesian inference that can scale to large datasets, allowing to sample from…

机器学习 · 统计学 2021-08-30 Mert Gürbüzbalaban , Xuefeng Gao , Yuanhan Hu , Lingjiong Zhu

Algorithmic stability is an important notion that has proven powerful for deriving generalization bounds for practical algorithms. The last decade has witnessed an increasing number of stability bounds for different algorithms applied on…

机器学习 · 统计学 2023-10-31 Lingjiong Zhu , Mert Gurbuzbalaban , Anant Raj , Umut Simsekli

While low-precision optimization has been widely used to accelerate deep learning, low-precision sampling remains largely unexplored. As a consequence, sampling is simply infeasible in many large-scale scenarios, despite providing…

机器学习 · 计算机科学 2022-06-22 Ruqi Zhang , Andrew Gordon Wilson , Christopher De Sa

Asynchronous stochastic gradient descent (SGD) enables scalable distributed training but suffers from gradient staleness. Existing mitigation strategies, such as delay-adaptive learning rates and staleness-aware filtering, typically…

机器学习 · 计算机科学 2026-05-15 Tehila Dahan , Roie Reshef , Sharon Goldstein , Kfir Y. Levy

Stochastic Gradient Descent with a constant learning rate (constant SGD) simulates a Markov chain with a stationary distribution. With this perspective, we derive several new results. (1) We show that constant SGD can be used as an…

机器学习 · 统计学 2018-01-23 Stephan Mandt , Matthew D. Hoffman , David M. Blei

In this paper we propose stochastic gradient-free methods and accelerated methods with momentum for solving stochastic optimization problems. All these methods rely on stochastic directions rather than stochastic gradients. We analyze the…

最优化与控制 · 数学 2020-01-15 Xiaopeng Luo , Xin Xu

We investigate the test risk of continuous-time stochastic gradient flow dynamics in learning theory. Using a path integral formulation we provide, in the regime of a small learning rate, a general formula for computing the difference…

机器学习 · 统计学 2025-03-05 Rodrigo Veiga , Anastasia Remizova , Nicolas Macris

Stochastic natural gradient variational inference (NGVI) is a popular and efficient algorithm for Bayesian inference. Despite empirical success, the convergence of this method is still not fully understood. In this work, we define and study…

统计方法学 · 统计学 2026-04-02 Thomas Guilmeau , Hadrien Hendrikx , Florence Forbes

Stochastic gradient descent is a classic algorithm that has gained great popularity especially in the last decades as the most common approach for training models in machine learning. While the algorithm has been well-studied when…

机器学习 · 统计学 2025-09-09 Jose Blanchet , Aleksandar Mijatović , Wenhao Yang

Obtaining coarse-grained models that accurately incorporate finite-size effects is an important open challenge in the study of complex, multi-scale systems. We apply Langevin regression, a recently developed method for finding stochastic…

适应与自组织系统 · 物理学 2021-10-12 Jordan Snyder , Jared L. Callaham , Steven L. Brunton , J. Nathan Kutz

We study a distributed consensus-based stochastic gradient descent (SGD) algorithm and show that the rate of convergence involves the spectral properties of two matrices: the standard spectral gap of a weight matrix from the network…

最优化与控制 · 数学 2016-09-02 Avleen S. Bijral , Anand D. Sarwate , Nathan Srebro
‹ 上一页 1 8 9 10 下一页 ›