English
Related papers

Related papers: Self-normalized Cram\'er-type Moderate Deviation o…

200 papers

We introduce a novel and efficient algorithm called the stochastic approximate gradient descent (SAGD), as an alternative to the stochastic gradient descent for cases where unbiased stochastic gradients cannot be trivially obtained.…

Machine Learning · Computer Science 2020-02-14 Yixuan Qiu , Xiao Wang

In this paper, we investigate the Milstein numerical scheme with step size $\eta$ for a stochastic differential equation driven by multiplicative Brownian motion. Under some appropriate coefficient conditions, the continuous-time system and…

Probability · Mathematics 2025-10-06 Peng Chen , Hui Jiang , Jing Wang

We propose a stochastic modified equations (SME) for modeling the asynchronous stochastic gradient descent (ASGD) algorithms. The resulting SME of Langevin type extracts more information about the ASGD dynamics and elucidates the…

Machine Learning · Statistics 2020-03-04 Jing An , Jianfeng Lu , Lexing Ying

One way to avoid overfitting in machine learning is to use model parameters distributed according to a Bayesian posterior given the data, rather than the maximum likelihood estimator. Stochastic gradient Langevin dynamics (SGLD) is one…

Machine Learning · Statistics 2017-12-05 Gaétan Marceau-Caron , Yann Ollivier

Applying standard Markov chain Monte Carlo (MCMC) algorithms to large data sets is computationally expensive. Both the calculation of the acceptance probability and the creation of informed proposals usually require an iteration through the…

Machine Learning · Statistics 2015-06-15 Yee Whye Teh , Alexandre Thiéry , Sebastian Vollmer

Let $(\xi_i,\mathcal{F}_i)_{i\geq1}$ be a sequence of martingale differences. Set $S_n=\sum_{i=1}^n\xi_i $ and $[ S]_n=\sum_{i=1}^n \xi_i^2.$ We prove a Cram\'er type moderate deviation expansion for $\mathbf{P}(S_n/\sqrt{[ S]_n} \geq x)$…

Probability · Mathematics 2020-05-11 Xiequan Fan , Ion Grama , Quansheng Liu , Qi-Man Shao

Stochastic gradient Langevin dynamics (SGLD) has gained the attention of optimization researchers due to its global optimization properties. This paper proves an improved convergence property to local minimizers of nonconvex objective…

Machine Learning · Computer Science 2024-07-08 Zhishen Huang , Stephen Becker

This paper considers mean square error (MSE) analysis for stochastic gradient sampling algorithms applied to underdamped Langevin dynamics under a global convexity assumption. A novel discrete Poisson equation framework is developed to…

Numerical Analysis · Mathematics 2025-11-07 Jianfeng Lu , Xuda Ye , Zhennan Zhou

Let $(g_{n})_{n\geq 1}$ be a sequence of independent and identically distributed positive random $d\times d$ matrices and consider the matrix product $G_n: = g_n \ldots g_1$. Under suitable conditions, we establish the Berry-Esseen bounds…

Probability · Mathematics 2020-10-02 Hui Xiao , Ion Grama , Quansheng Liu

In previous work, we introduced a method for determining convergence rates for integration methods for the kinetic Langevin equation for $M$-$\nabla$Lipschitz $m$-log-concave densities [arXiv:2302.10684, 2023]. In this article, we exploit…

Numerical Analysis · Mathematics 2023-06-16 Benedict Leimkuhler , Daniel Paulin , Peter A. Whalley

Recent years have seen advances in generalization bounds for noisy stochastic algorithms, especially stochastic gradient Langevin dynamics (SGLD) based on stability (Mou et al., 2018; Li et al., 2020) and information theoretic approaches…

Machine Learning · Computer Science 2022-11-02 Arindam Banerjee , Tiancong Chen , Xinyan Li , Yingxue Zhou

Let $X_1,X_2,...$ be independent random variables with zero means and finite variances, and let $S_n=\sum_{i=1}^nX_i$ and $V^2_n=\sum_{i=1}^nX^2_i$. A Cram\'{e}r type moderate deviation for the maximum of the self-normalized sums…

Statistics Theory · Mathematics 2013-07-24 Weidong Liu , Qi-Man Shao , Qiying Wang

The current interpretation of stochastic gradient descent (SGD) as a stochastic process lacks generality in that its numerical scheme restricts continuous-time dynamics as well as the loss function and the distribution of gradient noise. We…

Machine Learning · Statistics 2019-11-21 Soma Yokoi , Issei Sato

We establish Cram\'er-type moderate deviation theorems for sums of locally dependent random variables and combinatorial central limit theorems. Under some mild exponential moment conditions, optimal error bounds and convergence ranges are…

Probability · Mathematics 2021-12-22 Song-Hao Liu , Zhuo-Song Zhang

Stochastic gradient MCMC methods, such as stochastic gradient Langevin dynamics (SGLD), employ fast but noisy gradient estimates to enable large-scale posterior sampling. Although we can easily extend SGLD to distributed settings, it…

Machine Learning · Statistics 2021-06-16 Khaoula El Mekkaoui , Diego Mesquita , Paul Blomstedt , Samuel Kaski

Exponentiated gradient descent (EGD), a biologically motivated optimisation algorithm that respects Dale's law, produces log-normally distributed synaptic weights at convergence, in alignment with experimental observations in neuroscience.…

Machine Learning · Computer Science 2026-05-26 Nishanth Shetty , Madhava Prasath , Chandra Sekhar Seelamantula

We establish a Cram\'er-type moderate deviation theorem for double-index permutation statistics (DIPS). To the best of our knowledge, previous results only provided Berry-Esseen type bounds for DIPS, which cannot yield moderate deviation…

Probability · Mathematics 2026-03-27 Songhao Liu , Qiman Shao , Jingyu Xu

We provide a new convergence analysis of stochastic gradient Langevin dynamics (SGLD) for sampling from a class of distributions that can be non-log-concave. At the core of our approach is a novel conductance analysis of SGLD using an…

Machine Learning · Computer Science 2021-02-24 Difan Zou , Pan Xu , Quanquan Gu

We establish generalization error bounds for stochastic gradient Langevin dynamics (SGLD) with constant learning rate under the assumptions of dissipativity and smoothness, a setting that has received increased attention in the…

Machine Learning · Statistics 2021-11-29 Tyler Farghly , Patrick Rebeschini

We provide non-asymptotic convergence rates of the Polyak-Ruppert averaged stochastic gradient descent (SGD) to a normal random vector for a class of twice-differentiable test functions. A crucial intermediate step is proving a…

Statistics Theory · Mathematics 2019-04-04 Andreas Anastasiou , Krishnakumar Balasubramanian , Murat A. Erdogdu