English
Related papers

Related papers: Improved Finite-Particle Convergence Rates for Ste…

200 papers

Persistence diagrams (PDs) play a key role in topological data analysis (TDA), in which they are routinely used to describe topological properties of complicated shapes. PDs enjoy strong stability properties and have proven their utility in…

Computational Geometry · Computer Science 2017-11-10 Mathieu Carrière , Marco Cuturi , Steve Oudot

The subject of this paper is the estimation of a probability measure on ${\mathbb R}^d$ from data observed with an additive noise, under the Wasserstein metric of order $p$ (with $p\geq 1$). We assume that the distribution of the errors is…

Statistics Theory · Mathematics 2013-07-22 Jérôme Dedecker , Bertrand Michel

Sliced Stein discrepancy (SSD) and its kernelized variants have demonstrated promising successes in goodness-of-fit tests and model learning in high dimensions. Despite their theoretical elegance, their empirical performance depends…

Machine Learning · Computer Science 2021-07-22 Wenbo Gong , Kaibo Zhang , Yingzhen Li , José Miguel Hernández-Lobato

Wasserstein gradient flows are continuous time dynamics that define curves of steepest descent to minimize an objective function over the space of probability measures (i.e., the Wasserstein space). This objective is typically a divergence…

Optimization and Control · Mathematics 2021-02-23 Adil Salim , Anna Korba , Giulia Luise

The stochastic gradient descent (SGD) optimization algorithm plays a central role in a series of machine learning applications. The scientific literature provides a vast amount of upper error bounds for the SGD method. Much less attention…

Numerical Analysis · Mathematics 2020-10-05 Arnulf Jentzen , Philippe von Wurstemberger

The purpose of this paper is to answer a few open questions in the interface of kernel methods and PDE gradient flows. Motivated by recent advances in machine learning, particularly in generative modeling and sampling, we present a rigorous…

Machine Learning · Statistics 2024-10-29 Jia-Jie Zhu , Alexander Mielke

This work presents a multilevel variant of Stein variational gradient descent to more efficiently sample from target distributions. The key ingredient is a sequence of distributions with growing fidelity and costs that converges to the…

Numerical Analysis · Mathematics 2021-04-06 Terrence Alsup , Luca Venturi , Benjamin Peherstorfer

In this note, we establish a descent lemma for the population limit Mirrored Stein Variational Gradient Method~(MSVGD). This descent lemma does not rely on the path information of MSVGD but rather on a simple assumption for the mirrored…

Optimization and Control · Mathematics 2022-06-22 Lukang Sun , Peter Richtárik

We propose a novel approach to numerically approximate McKean-Vlasov stochastic differential equations (MV-SDE) using stochastic gradient descent (SGD) while avoiding the use of interacting particle systems (IPS) {and the associated…

Numerical Analysis · Mathematics 2026-01-22 Ankush Agarwal , Andrea Amato , Goncalo dos Reis , Stefano Pagliarani

We establish existence of Stein kernels for probability measures on $\mathbb{R}^d$ satisfying a Poincar\'e inequality, and obtain bounds on the Stein discrepancy of such measures. Applications to quantitative central limit theorems are…

Probability · Mathematics 2018-03-09 Thomas A. Courtade , Max Fathi , Ashwin Pananjady

We present and study a novel algorithm for the computation of 2-Wasserstein population barycenters of absolutely continuous probability measures on Euclidean space. The proposed method can be seen as a stochastic gradient descent procedure…

Optimization and Control · Mathematics 2023-10-24 Julio Backhoff-Veraguas , Joaquin Fontbona , Gonzalo Rios , Felipe Tobar

This study addresses the inverse problem of parameter estimation for Stochastic Differential Equations (SDEs) by minimizing a regularized discrepancy functional via Stochastic Gradient Descent (SGD). To achieve computational efficiency, we…

Machine Learning · Statistics 2026-03-31 Francisco Delgado-Vences , José Julián Pavón-Español , Arelly Ornelas

We establish subgeometric bounds on convergence rate of general Markov processes in the Wasserstein metric. In the discrete time setting we prove that the Lyapunov drift condition and the existence of a "good" $d$-small set imply…

Probability · Mathematics 2014-03-20 Oleg Butkovsky

Stochastic gradient descent (SGD) is a simple and popular method to solve stochastic optimization problems which arise in machine learning. For strongly convex problems, its convergence rate was known to be O(\log(T)/T), by running SGD for…

Machine Learning · Computer Science 2015-03-19 Alexander Rakhlin , Ohad Shamir , Karthik Sridharan

This paper proposes a novel parameter selection strategy for kernel-based gradient descent (KGD) algorithms, integrating bias-variance analysis with the splitting method. We introduce the concept of empirical effective dimension to quantify…

Machine Learning · Statistics 2026-03-05 Xiaotong Liu , Yunwen Lei , Xiangyu Chang , Shao-Bo Lin

The non-asymptotic analysis of Stochastic Gradient Descent (SGD) typically yields bounds that decompose into a bias term and a variance term. In this work, we focus on the bias component and study the extent to which SGD can match the…

Optimization and Control · Mathematics 2026-02-02 Daniel Cortild , Lucas Ketels , Juan Peypouquet , Guillaume Garrigos

Quantitative convergence in Wasserstein distance is often easier to establish than that in total variation distance. We show that such bounds allowing subgeometric rates yield central limit theorems (CLTs) for additive functionals of Markov…

Statistics Theory · Mathematics 2025-11-25 Rui Jin , Aixin Tan

Convergence rates of kernel density estimators for stationary time series are well studied. For invertible linear processes, we construct a new density estimator that converges, in the supremum norm, at the better, parametric, rate…

Statistics Theory · Mathematics 2009-09-29 Anton Schick , Wolfgang Wefelmeyer

We study the contraction in Wasserstein distance of the coordinate ascent variational inference algorithm. This is shown to hold under a transport-information inequality at the fixed points and a functional smoothness condition. The results…

Machine Learning · Statistics 2026-05-29 Rocco Caprio , Adrien Corenflos , Sam Power

We determine the convergence speed of a numerical scheme for approximating one-dimensional continuous strong Markov processes. The scheme is based on the construction of coin tossing Markov chains whose laws can be embedded into the process…

Probability · Mathematics 2020-08-26 Stefan Ankirchner , Thomas Kruse , Mikhail Urusov
‹ Prev 1 8 9 10 Next ›