Related papers: Self-normalized Cram\'er-type Moderate Deviation o…
We introduce a novel and efficient algorithm called the stochastic approximate gradient descent (SAGD), as an alternative to the stochastic gradient descent for cases where unbiased stochastic gradients cannot be trivially obtained.…
In this paper, we investigate the Milstein numerical scheme with step size $\eta$ for a stochastic differential equation driven by multiplicative Brownian motion. Under some appropriate coefficient conditions, the continuous-time system and…
We propose a stochastic modified equations (SME) for modeling the asynchronous stochastic gradient descent (ASGD) algorithms. The resulting SME of Langevin type extracts more information about the ASGD dynamics and elucidates the…
One way to avoid overfitting in machine learning is to use model parameters distributed according to a Bayesian posterior given the data, rather than the maximum likelihood estimator. Stochastic gradient Langevin dynamics (SGLD) is one…
Applying standard Markov chain Monte Carlo (MCMC) algorithms to large data sets is computationally expensive. Both the calculation of the acceptance probability and the creation of informed proposals usually require an iteration through the…
Let $(\xi_i,\mathcal{F}_i)_{i\geq1}$ be a sequence of martingale differences. Set $S_n=\sum_{i=1}^n\xi_i $ and $[ S]_n=\sum_{i=1}^n \xi_i^2.$ We prove a Cram\'er type moderate deviation expansion for $\mathbf{P}(S_n/\sqrt{[ S]_n} \geq x)$…
Stochastic gradient Langevin dynamics (SGLD) has gained the attention of optimization researchers due to its global optimization properties. This paper proves an improved convergence property to local minimizers of nonconvex objective…
This paper considers mean square error (MSE) analysis for stochastic gradient sampling algorithms applied to underdamped Langevin dynamics under a global convexity assumption. A novel discrete Poisson equation framework is developed to…
Let $(g_{n})_{n\geq 1}$ be a sequence of independent and identically distributed positive random $d\times d$ matrices and consider the matrix product $G_n: = g_n \ldots g_1$. Under suitable conditions, we establish the Berry-Esseen bounds…
In previous work, we introduced a method for determining convergence rates for integration methods for the kinetic Langevin equation for $M$-$\nabla$Lipschitz $m$-log-concave densities [arXiv:2302.10684, 2023]. In this article, we exploit…
Recent years have seen advances in generalization bounds for noisy stochastic algorithms, especially stochastic gradient Langevin dynamics (SGLD) based on stability (Mou et al., 2018; Li et al., 2020) and information theoretic approaches…
Let $X_1,X_2,...$ be independent random variables with zero means and finite variances, and let $S_n=\sum_{i=1}^nX_i$ and $V^2_n=\sum_{i=1}^nX^2_i$. A Cram\'{e}r type moderate deviation for the maximum of the self-normalized sums…
The current interpretation of stochastic gradient descent (SGD) as a stochastic process lacks generality in that its numerical scheme restricts continuous-time dynamics as well as the loss function and the distribution of gradient noise. We…
We establish Cram\'er-type moderate deviation theorems for sums of locally dependent random variables and combinatorial central limit theorems. Under some mild exponential moment conditions, optimal error bounds and convergence ranges are…
Stochastic gradient MCMC methods, such as stochastic gradient Langevin dynamics (SGLD), employ fast but noisy gradient estimates to enable large-scale posterior sampling. Although we can easily extend SGLD to distributed settings, it…
Exponentiated gradient descent (EGD), a biologically motivated optimisation algorithm that respects Dale's law, produces log-normally distributed synaptic weights at convergence, in alignment with experimental observations in neuroscience.…
We establish a Cram\'er-type moderate deviation theorem for double-index permutation statistics (DIPS). To the best of our knowledge, previous results only provided Berry-Esseen type bounds for DIPS, which cannot yield moderate deviation…
We provide a new convergence analysis of stochastic gradient Langevin dynamics (SGLD) for sampling from a class of distributions that can be non-log-concave. At the core of our approach is a novel conductance analysis of SGLD using an…
We establish generalization error bounds for stochastic gradient Langevin dynamics (SGLD) with constant learning rate under the assumptions of dissipativity and smoothness, a setting that has received increased attention in the…
We provide non-asymptotic convergence rates of the Polyak-Ruppert averaged stochastic gradient descent (SGD) to a normal random vector for a class of twice-differentiable test functions. A crucial intermediate step is proving a…