Related papers: Wasserstein Gradient Flows of the Discrepancy with…
Maximum mean discrepancy (MMD) flows suffer from high computational costs in large scale computations. In this paper, we show that MMD flows with Riesz kernels $K(x,y) = - \|x-y\|^r$, $r \in (0,2)$ have exceptional properties which allow…
We study geometric properties of the gradient flow for learning deep linear convolutional networks. For linear fully connected networks, it has been shown recently that the corresponding gradient flow on parameter space can be written as a…
We prove the equivalence between the notion of Wasserstein gradient flow for a one-dimensional nonlocal transport PDE with attractive/repulsive Newtonian potential on one side, and the notion of entropy solution of a Burgers-type scalar…
Motivated by a constrained minimization problem, it is studied the gradient flows with respect to Hessian Riemannian metrics induced by convex functions of Legendre type. The first result characterizes Hessian Riemannian structures on…
We study properties of an attractive-repulsive energy functional based on power-kernels, which can be used for halftoning of images. In the first part of this work, using a variational framework for probability measures, we examine…
We define new differential structures on the Wasserstein spaces $\mathcal{W}_p(M)$ for $p > 2$ and a general Riemannian manifold $(M,g)$. We consider a very general and possibly degenerate second order partial differential flow equation…
In this work, we investigate links between the formulation of the flow of marginals of reversible diffusion processes as gradient flows in the space of probability measures and path wise large deviation principles for sequences of such…
Wasserstein gradient flow (WGF) is a common method to perform optimization over the space of probability measures. While WGF is guaranteed to converge to a first-order stationary point, for nonconvex functionals the converged solution does…
In this paper, we address the classification of instances each characterized not by a singular point, but by a distribution on a vector space. We employ the Wasserstein metric to measure distances between distributions, which are then used…
This work characterizes, analytically and numerically, two major effects of the quadratic Wasserstein ($W_2$) distance as the measure of data discrepancy in computational solutions of inverse problems. First, we show, in the…
We use Stein's method to bound the Wasserstein distance of order $2$ between a measure $\nu$ and the Gaussian measure using a stochastic process $(X_t)_{t \geq 0}$ such that $X_t$ is drawn from $\nu$ for any $t > 0$. If the stochastic…
The Wasserstein distances $W_p$ ($p\geq 1$), defined in terms of solution to the Monge-Kantorovich problem, are known to be a useful tool to investigate transport equations. In particular, the Benamou-Brenier formula characterizes the…
We investigate long-time behaviors of empirical measures associated with subordinated Dirichlet diffusion processes on a compact Riemannian manifold $M$ with boundary $\partial M$ to some reference measure, under the quadratic Wasserstein…
We study rays and co-rays in the Wasserstein space $P_p(\mathcal{X})$ ($p > 1$) whose ambient space $\mathcal{X}$ is a complete, separable, non-compact, locally compact length space. We show that rays in the Wasserstein space can be…
The sliced-Wasserstein flow is an evolution equation where a probability density evolves in time, advected by a velocity field computed as the average among directions in the unit sphere of the optimal transport displacements from its 1D…
We develop a framework for generalized variational inference in infinite-dimensional function spaces and use it to construct a method termed Gaussian Wasserstein inference (GWI). GWI leverages the Wasserstein distance between Gaussian…
This paper studies convergence of empirical measures smoothed by a Gaussian kernel. Specifically, consider approximating $P\ast\mathcal{N}_\sigma$, for $\mathcal{N}_\sigma\triangleq\mathcal{N}(0,\sigma^2 \mathrm{I}_d)$, by…
Stacking Gaussian Processes severely diminishes the model's ability to detect outliers, which when combined with non-zero mean functions, further extrapolates low non-parametric variance to low training data density regions. We propose a…
The defining equation $(\ast):\ \dot \omega\_t=-F'(\omega\_t),$ of a gradient flow is kinetic in essence. This article explores some dynamical (rather than kinetic) features of gradient flows (i) by embedding equation $(\ast)$ into the…
Despite the widespread success of Transformers across various domains, their optimization guarantees in large-scale model settings are not well-understood. This paper rigorously analyzes the convergence properties of gradient flow in…