English
Related papers

Related papers: Theoretical guarantees for stochastic gradient sam…

200 papers

Stochastic gradient Langevin dynamics (SGLD) is a computationally efficient sampler for Bayesian posterior inference given a large scale dataset. Although SGLD is designed for unbounded random variables, many practical models incorporate…

Machine Learning · Statistics 2019-06-21 Soma Yokoi , Takuma Otsuka , Issei Sato

This work introduces a general framework for establishing the long time accuracy for approximations of Markovian dynamical systems on separable Banach spaces. Our results illuminate the role that a certain uniformity in Wasserstein…

Numerical Analysis · Mathematics 2023-02-06 Nathan E. Glatt-Holtz , Cecilia F. Mondaini

Sliced Wasserstein distances preserve properties of classic Wasserstein distances while being more scalable for computation and estimation in high dimensions. The goal of this work is to quantify this scalability from three key aspects: (i)…

Machine Learning · Statistics 2022-10-18 Sloan Nietert , Ritwik Sadhu , Ziv Goldfeld , Kengo Kato

This paper presents a new approach to the classical problem of quantifying posterior contraction rates (PCRs) in Bayesian statistics. Our approach relies on Wasserstein distance, and it leads to two main contributions which improve on the…

Statistics Theory · Mathematics 2022-05-03 Emanuele Dolera , Stefano Favaro , Edoardo Mainini

Stochastic Gradient Langevin Dynamics (SGLD) is a sampling scheme for Bayesian modeling adapted to large datasets and models. SGLD relies on the injection of Gaussian Noise at each step of a Stochastic Gradient Descent (SGD) update. In this…

Machine Learning · Computer Science 2018-06-11 Henri Palacci , Henry Hess

Gradient clipping is a popular modification to standard (stochastic) gradient descent, at every iteration limiting the gradient norm to a certain value $c >0$. It is widely used for example for stabilizing the training of deep learning…

Machine Learning · Computer Science 2023-11-10 Anastasia Koloskova , Hadrien Hendrikx , Sebastian U. Stich

Stochastic Gradient (SG) Markov Chain Monte Carlo algorithms (MCMC) are popular algorithms for Bayesian sampling in the presence of large datasets. However, they come with little theoretical guarantees and assessing their empirical…

Machine Learning · Statistics 2024-05-16 Lorenzo Mauri , Giacomo Zanella

Self-supervised learning is one of the most promising approaches to acquiring knowledge from limited labeled data. Despite the substantial advancements made in recent years, self-supervised models have posed a challenge to practitioners, as…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Franciskus Xaverius Erick , Mina Rezaei , Johanna Paula Müller , Bernhard Kainz

The existence and uniqueness of the numerical invariant measure of the backward Euler-Maruyama method for stochastic differential equations with Markovian switching is yielded, and it is revealed that the numerical invariant measure…

Probability · Mathematics 2022-11-04 Xiaoyue Li , Qianlin Ma , Hongfu Yang , Chenggui Yuan

Perturbation theory for Markov chains addresses the question how small differences in the transitions of Markov chains are reflected in differences between their distributions. We prove powerful and flexible bounds on the distance of the…

Computation · Statistics 2017-02-27 Daniel Rudolf , Nikolaus Schweizer

Using entropic inequalities from information theory, we provide new bounds on the total variation and 2-Wasserstein distances between a conditionally Gaussian law and a Gaussian law with invertible covariance matrix. We apply our results to…

Probability · Mathematics 2025-06-04 Lucia Celli , Giovanni Peccati

The space of Gaussian measures on a Euclidean space is geodesically convex in the $L^2$-Wasserstein space. This space is a finite dimensional manifold since Gaussian measures are parameterized by means and covariance matrices. By…

Differential Geometry · Mathematics 2009-02-11 Asuka Takatsu

In this article, we study the problem of sampling from distributions whose densities are not necessarily smooth nor logconcave. We propose a simple Langevin-based algorithm that does not rely on popular but computationally challenging…

Machine Learning · Statistics 2025-12-02 Tim Johnston , Iosif Lytras , Nikolaos Makras , Sotirios Sabanis

Gromov--Wasserstein (GW) distances compare graphs, shapes, and point clouds through internal distances, without requiring a common coordinate system. This invariance is powerful, but discrete GW is a nonconvex quadratic optimal transport…

Machine Learning · Computer Science 2026-05-15 Ao Xu , Tieru Wu

We propose a fully discrete variational scheme for nonlinear evolution equations with gradient flow structure on the space of finite Radon measures on an interval with respect to a generalized version of the Wasserstein distance with…

Numerical Analysis · Mathematics 2016-09-29 Jonathan Zinsl , Daniel Matthes

Stochastic gradient descent (SGD) and its variants have established themselves as the go-to algorithms for large-scale machine learning problems with independent samples due to their generalization performance and intrinsic computational…

Machine Learning · Statistics 2025-08-25 Hao Chen , Lili Zheng , Raed Al Kontar , Garvesh Raskutti

We revisit Markowitz's mean-variance portfolio selection model by considering a distributionally robust version, where the region of distributional uncertainty is around the empirical measure and the discrepancy between probability measures…

Methodology · Statistics 2018-02-15 Jose Blanchet , Lin Chen , Xun Yu Zhou

We consider the long time statistics of a one-dimensional stochastic Ginzburg-Landau equation with cubic nonlinearity while being subjected to random perturbations via an additive Gaussian noise. Under the assumption that sufficiently many…

Probability · Mathematics 2024-05-21 Hung D. Nguyen

This paper presents a novel Wasserstein distributionally robust control and state estimation algorithm for partially observable linear stochastic systems, where the probability distributions of disturbances and measurement noises are…

Systems and Control · Electrical Eng. & Systems 2024-06-05 Minhyuk Jang , Astghik Hakobyan , Insoon Yang

Optimal transport is a foundational problem in optimization, that allows to compare probability distributions while taking into account geometric aspects. Its optimal objective value, the Wasserstein distance, provides an important loss…

Machine Learning · Computer Science 2020-02-21 Marin Ballu , Quentin Berthet , Francis Bach