Related papers: A weak convergence approach to large deviations fo…
This paper focuses on systems of nonlinear second-order stochastic differential equations with multi-scales. The motivation for our study stems from mathematical physics and statistical mechanics, for examples, Langevin dynamics and…
The work concerns deviation estimates for multivalued McKean-Vlasov stochastic differential equations. First of all, we prove the large deviation principle for them by the weak convergence approach. Then the central limit theorem for them…
This study focuses on large deviation principles for fully coupled multiscale multivalued stochastic systems, in which the slow component is governed by a multivalued stochastic differential equation and the fast component is described by a…
This article suggests that deterministic Gradient Descent, which does not use any stochastic gradient approximation, can still exhibit stochastic behaviors. In particular, it shows that if the objective function exhibit multiscale…
The tuning of stochastic gradient algorithms (SGAs) for optimization and sampling is often based on heuristics and trial-and-error rather than generalizable theory. We address this theory--practice gap by characterizing the large-sample…
Large-scale optimization problems require algorithms both effective and efficient. One such popular and proven algorithm is Stochastic Gradient Descent which uses first-order gradient information to solve these problems. This paper studies…
We consider potential type dynamical systems in finite dimensions with two meta-stable states. They are subject to two sources of perturbation: a slow external periodic perturbation of period $T$ and a small Gaussian random perturbation of…
In this paper we establish the large deviation principle for the stochastic quasi-geostrophic equation in the subcritical case with small multiplicative noise. The proof is mainly based on the stochastic control and weak convergence…
We propose a computational method for large deviation statistics of time-averaged quantities in general Markov processes. In our proposed method, we repeat a response measurement against external forces, where the forces are determined by…
We study large deviations in the context of stochastic gradient descent for one-hidden-layer neural networks with quadratic loss. We derive a quenched large deviation principle, where we condition on an initial weight measure, and an…
This paper is concerned with the large deviation principle of the non-local fractional stochastic reaction-diffusion equation with a polynomial drift of arbitrary degree driven by multiplicative noise defined on unbounded domains. We first…
Motivated by learning of correlated equilibria in non-cooperative games, we perform a large deviations analysis of a regret minimizing stochastic approximation algorithm. The regret minimization algorithm we consider comprises multiple…
Using the weak convergence approach, we prove the large deviation principle (LDP) for solutions to quasilinear stochastic evolution equations with small Gaussian noise in the critical variational setting, a recently developed general…
The Freidlin-Wentzell large deviation principle is established for the distributions of stochastic evolution equations with general monotone drift and small multiplicative noise. As examples, the main results are applied to derive the large…
Many machine learning and optimization algorithms are built upon the framework of stochastic approximation (SA), for which the selection of step-size (or learning rate) $\{\alpha_n\}$ is crucial for success. An essential condition for…
This paper establishes the first almost sure convergence rate and the first maximal concentration bound with exponential tails for general contractive stochastic approximation algorithms with Markovian noise. As a corollary, we also obtain…
The ODE method has been a workhorse for algorithm design and analysis since the introduction of the stochastic approximation. It is now understood that convergence theory amounts to establishing robustness of Euler approximations for ODEs,…
When training neural networks, it has been widely observed that a large step size is essential in stochastic gradient descent (SGD) for obtaining superior models. However, the effect of large step sizes on the success of SGD is not well…
In this paper, we establish a large deviation principle for the stochastic generalized Ginzburg-Landau equation driven by jump noise. The main difficulties come from the highly non-linear coefficient. Here we adopt a new sufficient…
Stochastic gradient descent is one of the most successful approaches for solving large-scale problems, especially in machine learning and statistics. At each iteration, it employs an unbiased estimator of the full gradient computed from one…