Related papers: On the stability of the stochastic gradient Langev…
This article suggests that deterministic Gradient Descent, which does not use any stochastic gradient approximation, can still exhibit stochastic behaviors. In particular, it shows that if the objective function exhibit multiscale…
Stochastic-approximation gradient methods are attractive for large-scale convex optimization because they offer inexpensive iterations. They are especially popular in data-fitting and machine-learning applications where the data arrives in…
This paper considers the problem of learning, from samples, the dependency structure of a system of linear stochastic differential equations, when some of the variables are latent. In particular, we observe the time evolution of some…
Sampling from discrete distributions is a ubiquitous task in machine learning, recently revisited by the emergence of discrete diffusion models. While Langevin algorithms constitute the state of the art for continuous spaces, discrete…
We provide a framework to analyze the convergence of discretized kinetic Langevin dynamics for $M$-$\nabla$Lipschitz, $m$-convex potentials. Our approach gives convergence rates of $\mathcal{O}(m/M)$, with explicit stepsize restrictions,…
A wide body of work has applied the concept of critical slowing down to estimate the stability of different Earth system components. Most of them -- such as global vegetation -- are inherently non-stationary, for example due to strong…
In this paper we address the problem of consistently construct Langevin equations to describe fluctuations in non-linear systems. Detailed balance severely restricts the choice of the random force, but we prove that this property together…
Non-Markovian stochastic Langevin-like equations of motion are compared to their corresponding Markovian (local) approximations. The validity of the local approximation for these equations, when contrasted with the fully nonlocal ones, is…
We employ constraints to control the parameter space of deep neural networks throughout training. The use of customized, appropriately designed constraints can reduce the vanishing/exploding gradients problem, improve smoothness of…
The so-called SAGA-LD algorithm is used for efficient sampling from high-dimensional distributions in machine learning. Its intricate dynamics resists standard approaches of Markov chain theory. We prove, using a model-specific method, that…
We rigorously show that a local spin system giving rise to a slow Hamiltonian dynamics is stable against generic, even time-dependent, local perturbations. The sum of these perturbations can cover a significant amount of the system's size.…
We describe a criterion for particles suspended in a randomly moving fluid to aggregate. Aggregation occurs when the expectation value of a random variable is negative. This random variable evolves under a stochastic differential equation.…
The classical convergence analysis of SGD is carried out under the assumption that the norm of the stochastic gradient is uniformly bounded. While this might hold for some loss functions, it is violated for cases where the objective…
We revisit processes generated by iterated random functions driven by a stationary and ergodic sequence. Such a process is called strongly stable if a random initialization exists, for which the process is stationary and ergodic, and for…
We consider variational inequality solutions with prescribed gradient constraints for first order linear boundary value problems. For operators with coefficients only in $L^2$, we show the existence and uniqueness of the solution by using a…
Constrained sampling is an important and challenging task in computational statistics, concerned with generating samples from a distribution under certain constraints. There are numerous types of algorithm aimed at this task, ranging from…
This note provides a simple derivation of the overdamped approximation for kinetic (or underdamped) equilibrium Langevin dynamics, in cases where certain coefficients depend on the position variable. The equivalent small-mass limit of these…
In this paper, we prove convergence in distribution of Langevin processes in the overdamped asymptotics. The proof relies on the classical perturbed test function (or corrector) method, which is used both to show tightness in path space,…
We study to what extent may stochastic gradient descent (SGD) be understood as a "conventional" learning rule that achieves generalization performance by obtaining a good fit to training data. We consider the fundamental stochastic convex…
Linear response theory is a fundamental framework studying the macroscopic response of a physical system to an external perturbation. This paper focuses on the rigorous mathematical justification of linear response theory for Langevin…