English
Related papers

Related papers: Straight-Through Estimator as Projected Wasserstei…

200 papers

Stream stochastic gradient descent (SGD) is a simple and efficient method for solving online optimization problems in operations research (OR), where data is generated by parameter-dependent Markov chains. Unlike traditional approaches…

Optimization and Control · Mathematics 2025-09-03 Xiang Li , Jiadong Liang , Xinyun Chen , Zhihua Zhang

For statistical models on circles, we investigate performance of estimators defined as the projections of the empirical distribution with respect to the Wasserstein distance. We develop algorithms for computing the Wasserstein projection…

Statistics Theory · Mathematics 2025-10-22 Naoki Otani , Takeru Matsuda

Generalized sliced Wasserstein distance is a variant of sliced Wasserstein distance that exploits the power of non-linear projection through a given defining function to better capture the complex structures of the probability…

Machine Learning · Statistics 2022-10-20 Dung Le , Huy Nguyen , Khai Nguyen , Trang Nguyen , Nhat Ho

The Sliced-Wasserstein distance (SW) is being increasingly used in machine learning applications as an alternative to the Wasserstein distance and offers significant computational and statistical benefits. Since it is defined as an…

Machine Learning · Statistics 2022-01-05 Kimia Nadjahi , Alain Durmus , Pierre E. Jacob , Roland Badeau , Umut Şimşekli

Wasserstein-Fisher-Rao (WFR) gradient flows have been recently proposed as a powerful sampling tool that combines the advantages of pure Wasserstein (W) and pure Fisher-Rao (FR) gradient flows. Existing algorithmic developments implicitly…

Machine Learning · Statistics 2026-03-02 Francesca Romana Crucinio , Sahani Pathiraja

By the continuous mapping theorem, if a sequence of $d$-dimensional random vectors $(\mathbf{W}_n)_{n\geq1}$ converges in distribution to a multivariate normal random variable $\Sigma^{1/2}\mathbf{Z}$, then the sequence of random variables…

Probability · Mathematics 2020-03-18 Robert E. Gaunt

We study the Wasserstein natural gradient in parametric statistical models with continuous sample spaces. Our approach is to pull back the $L^2$-Wasserstein metric tensor in the probability density space to a parameter space, equipping the…

Optimization and Control · Mathematics 2024-08-20 Yifan Chen , Wuchen Li

The sliced Wasserstein (SW) distances between two probability measures are defined as the expectation of the Wasserstein distance between two one-dimensional projections of the two measures. The randomness comes from a projecting direction…

Machine Learning · Statistics 2024-02-20 Khai Nguyen , Nhat Ho

We study a natural Wasserstein gradient flow on manifolds of probability distributions with discrete sample spaces. We derive the Riemannian structure for the probability simplex from the dynamical formulation of the Wasserstein distance on…

Optimization and Control · Mathematics 2021-04-19 Wuchen Li , Guido Montufar

Recently, optimization on the Riemannian manifold have provided valuable insights to the optimization community. In this regard, extending these methods to to the Wasserstein space is of particular interest, since optimization on…

Machine Learning · Computer Science 2025-11-05 Mingyang Yi , Bohan Wang

Wasserstein distances are metrics on probability distributions inspired by the problem of optimal mass transportation. Roughly speaking, they measure the minimal effort required to reconfigure the probability mass of one distribution in…

Methodology · Statistics 2019-04-10 Victor M. Panaretos , Yoav Zemel

The sliced Wasserstein flow (SWF), a nonparametric and implicit generative gradient flow, is transformed into a Liouville partial differential equation (PDE)-based formalism. First, the stochastic diffusive term from the Fokker-Planck…

Machine Learning · Statistics 2026-05-12 Jayshawn Cooper , Pilhwa Lee

Optimal Transport has sparked vivid interest in recent years, in particular thanks to the Wasserstein distance, which provides a geometrically sensible and intuitive way of comparing probability measures. For computational reasons, the…

Machine Learning · Computer Science 2024-03-19 Eloi Tanguy

Full waveform inversion (FWI) enables us to obtain high-resolution velocity models of the subsurface. However, estimating the associated uncertainties in the process is not trivial. Commonly, uncertainty estimation is performed within the…

Geophysics · Physics 2023-05-16 Muhammad Izzatullah , Matteo Ravasi , Tariq Alkhalifah

Predicting molecular ground-state conformation (i.e., energy-minimized conformation) is crucial for many chemical applications such as molecular docking and property prediction. Classic energy-based simulation is time-consuming when solving…

Biomolecules · Quantitative Biology 2025-05-23 Fanmeng Wang , Minjie Cheng , Hongteng Xu

Ranking distributions according to a stochastic order has wide applications in diverse areas. Although stochastic dominance has received much attention, convex order, particularly in general dimensions, has yet to be investigated from a…

Methodology · Statistics 2025-01-15 Jakwang Kim , Young-Heon Kim , Yuanlong Ruan , Andrew Warren

Stein variational gradient descent (SVGD) is a kernel-based particle method for sampling from a target distribution, e.g., in generative modeling and Bayesian inference. SVGD does not require estimating the gradient of the log-density,…

Machine Learning · Statistics 2025-04-10 Viktor Stein , Wuchen Li

Many applications in machine learning involve data represented as probability distributions. The emergence of such data requires radically novel techniques to design tractable gradient flows on probability distributions over this type of…

Machine Learning · Computer Science 2025-06-10 Clément Bonet , Christophe Vauthier , Anna Korba

We derive Wasserstein distance bounds between the probability distributions of a stochastic integral (It\^o) process with jumps $(X_t)_{t\in [0,T]}$ and a jump-diffusion process $(X^\ast_t)_{t\in [0,T]}$. Our bounds are expressed using the…

Probability · Mathematics 2022-12-12 Jean-Christophe Breton , Nicolas Privault

Wasserstein barycenters provide a principled approach for aggregating probability measures, while preserving the geometry of their ambient space. Existing discrete methods are not scalable as they assume access to the complete set of…

Machine Learning · Statistics 2026-03-10 Eduardo Fernandes Montesuma , Yassir Bendou , Mike Gartrell