English
Related papers

Related papers: Natural gradient via optimal transport

200 papers

Natural gradient descent is a principled method for adapting the parameters of a statistical model on-line using an underlying Riemannian parameter space to redefine the direction of steepest descent. The algorithm is examined via methods…

Disordered Systems and Neural Networks · Physics 2009-10-31 Magnus Rattray , David Saad

Bayesian inference plays an important role in advancing machine learning, but faces computational challenges when applied to complex models such as deep neural networks. Variational inference circumvents these challenges by formulating…

Machine Learning · Statistics 2018-08-03 Mohammad Emtiyaz Khan , Didrik Nielsen

Wasserstein geometry and information geometry are two important structures introduced in a manifold of probability distributions. The former is defined by using the transportation cost between two distributions, so it reflects the metric…

Statistics Theory · Mathematics 2020-03-13 Shun-ichi Amari

We study discretizations of Hamiltonian systems on the probability density manifold equipped with the $L^2$-Wasserstein metric. Based on discrete optimal transport theory, several Hamiltonian systems on graph (lattice) with different…

Numerical Analysis · Mathematics 2020-06-17 Jianbo Cui , Luca Dieci , Haomin Zhou

This paper reviews different numerical methods for specific examples of Wasserstein gradient flows: we focus on nonlinear Fokker-Planck equations,but also discuss discretizations of the parabolic-elliptic Keller-Segel model and of the…

Numerical Analysis · Mathematics 2020-03-10 Jose A. Carrillo , Daniel Matthes , Marie-Therese Wolfram

The Wasserstein probability metric has received much attention from the machine learning community. Unlike the Kullback-Leibler divergence, which strictly measures change in probability, the Wasserstein metric reflects the underlying…

Machine Learning · Computer Science 2017-06-01 Marc G. Bellemare , Ivo Danihelka , Will Dabney , Shakir Mohamed , Balaji Lakshminarayanan , Stephan Hoyer , Rémi Munos

We study the Wasserstein gradient flow of semi-discrete energies in the space of probability measures, that is functionals depending on two measures-one being an absolutely continuous density and the other an atomic measure. These energies…

Analysis of PDEs · Mathematics 2026-03-05 Joao Miguel Machado

Natural gradients have been widely used in optimization of loss functionals over probability space, with important examples such as Fisher-Rao gradient descent for Kullback-Leibler divergence, Wasserstein gradient descent for…

Numerical Analysis · Mathematics 2020-06-30 Lexing Ying

Modeling observations as random distributions embedded within Wasserstein spaces is becoming increasingly popular across scientific fields, as it captures the variability and geometric structure of the data more effectively. However, the…

Statistics Theory · Mathematics 2026-04-08 François Bachoc , Alberto González-Sanz , Jean-Michel Loubes , Yisha Yao

We introduce dynamic and static formulations that formally extend unbalanced optimal transport from the space of positive densities to the space of Riemannian metrics. The first construction is based on a dynamic variational formulation in…

Differential Geometry · Mathematics 2026-05-27 Martin Bauer , Peter W. Michor , François-Xavier Vialard

We study information matrices for statistical models by the $L^2$-Wasserstein metric. We call them Wasserstein information matrices (WIMs), which are analogs of classical Fisher information matrices. We introduce Wasserstein score functions…

Statistics Theory · Mathematics 2020-08-12 Wuchen Li , Jiaxi Zhao

We study a variant of the dynamical optimal transport problem in which the energy to be minimised is modulated by the covariance matrix of the distribution. Such transport metrics arise naturally in mean-field limits of certain ensemble…

Analysis of PDEs · Mathematics 2024-12-23 Martin Burger , Matthias Erbar , Franca Hoffmann , Daniel Matthes , André Schlichting

In this paper we establish a rigorous gradient flow structure for one-dimensional Kimura equations with respect to some Wasserstein-Shahshahani optimal transport geometry. This is achieved by first conditioning the underlying stochastic…

Analysis of PDEs · Mathematics 2022-10-03 Jean-Baptiste Casteras , Léonard Monsaingeon

We study the natural gradient method for learning in deep Bayesian networks, including neural networks. There are two natural geometries associated with such learning systems consisting of visible and hidden units. One geometry is related…

Machine Learning · Computer Science 2020-05-22 Nihat Ay

Optimal transport (OT) provides powerful tools for comparing probability measures in various types. The Wasserstein distance which arises naturally from the idea of OT is widely used in many machine learning applications. Unfortunately,…

Optimization and Control · Mathematics 2021-06-03 Shu Liu , Haodong Sun , Hongyuan Zha

We propose a fully discrete variational scheme for nonlinear evolution equations with gradient flow structure on the space of finite Radon measures on an interval with respect to a generalized version of the Wasserstein distance with…

Numerical Analysis · Mathematics 2016-09-29 Jonathan Zinsl , Daniel Matthes

Gaussian mixture models form a flexible and expressive parametric family of distributions that has found applications in a wide variety of applications. Unfortunately, fitting these models to data is a notoriously hard problem from a…

Statistics Theory · Mathematics 2023-01-05 Yuling Yan , Kaizheng Wang , Philippe Rigollet

We introduce an optimal transport topology on the space of probability measures over a fiber bundle, which penalizes the transport cost from one fiber to another. For simplicity, we illustrate our construction in the Euclidean case…

Analysis of PDEs · Mathematics 2024-01-12 Jan Peszek , David Poyato

Optimal transport has recently proved to be a useful tool in various machine learning applications needing comparisons of probability measures. Among these, applications of distributionally robust optimization naturally involve Wasserstein…

Optimization and Control · Mathematics 2023-03-24 Waïss Azizian , Franck Iutzeler , Jérôme Malick

Wasserstein distance (WD) and the associated optimal transport plan have been proven useful in many applications where probability measures are at stake. In this paper, we propose a new proxy of the squared WD, coined min-SWGG, that is…

Machine Learning · Statistics 2023-10-31 Guillaume Mahey , Laetitia Chapel , Gilles Gasso , Clément Bonet , Nicolas Courty