Related papers: Pinsker's inequality for adapted total variation
Optimal transport (OT) provides powerful tools for comparing probability measures in various types. The Wasserstein distance which arises naturally from the idea of OT is widely used in many machine learning applications. Unfortunately,…
Let $(W,H,\mu)$ be the classical Wiener space, assume that $U=I_W+u$ is an adapted perturbation of identity where the perturbation $u$ is an equivalence class w.r.to the Wiener measure. We study several necessary and sufficient conditions…
Birkhoff's Theorem states that doubly stochastic matrices are convex combinations of permutation matrices. Quantum mechanically these matrices are doubly stochastic channels, i.e. they are completely positive maps preserving both the trace…
The following type exponential convergence is proved for (non-degenerate or degenerate) McKean-Vlasov SDEs: $$W_2(\mu_t,\mu_\infty)^2 +{\rm Ent}(\mu_t|\mu_\infty)\le c {\rm e}^{-\lambda t} \min\big\{W_2(\mu_0, \mu_\infty)^2,{\rm…
A growing number of generative statistical models do not permit the numerical evaluation of their likelihood functions. Approximate Bayesian computation (ABC) has become a popular approach to overcome this issue, in which one simulates…
In this paper, we provide sufficient conditions for the existence of the invariant distribution and for subgeometric rates of convergence in Wasserstein distance for general state-space Markov chains which are (possibly) not irreducible.…
We discuss a relation between the Kantorovich-Wasserstein (KW) metric and the Kullback-Leibler (KL) divergence. The former is defined using the optimal transport problem (OTP) in the Kantorovich formulation. The latter is used to define…
The sliced Wasserstein distance as well as its variants have been widely considered in comparing probability measures defined on $\mathbb R^d$. Here we derive the notion of sliced Wasserstein distance for measures on an infinite dimensional…
The change in the normal between any two nearby points on a closed, smooth surface is bounded with respect to the local feature size (distance to the medial axis). An incorrect proof of this lemma appeared as part of the analysis of the…
We study optimization problems whereby the optimization variable is a probability measure. Since the probability space is not a vector space, many classical and powerful methods for optimization (e.g., gradients) are of little help. Thus,…
Correctly estimating the discrepancy between two data distributions has always been an important task in Machine Learning. Recently, Cuturi proposed the Sinkhorn distance which makes use of an approximate Optimal Transport cost between two…
We study distributionally robust optimization with Sinkhorn distance -- a variant of Wasserstein distance based on entropic regularization. We derive a convex programming dual reformulation for general nominal distributions, transport…
Let $\boldsymbol u$ be the minimizer of vectorial total variation ($VTV$) with $L^2$ data-fidelity term on an interval $I$. We show that the total variation of $\boldsymbol u$ over any subinterval of $I$ is bounded by that of the datum over…
In Part 1, we developed a new technique based on Lipschitz pushforwards for proving the jump set containment property $\mathcal{H}^{m-1}(J_u \setminus J_f)=0$ of solutions $u$ to total variation denoising. We demonstrated that the technique…
Wasserstein distances are metrics on probability distributions inspired by the problem of optimal mass transportation. Roughly speaking, they measure the minimal effort required to reconfigure the probability mass of one distribution in…
We derive first-order (in the stepsize) bounds on the bias in Wasserstein distances of the invariant measure of stochastic gradient kinetic Langevin dynamics with minimal assumptions on the stochastic gradient noise. These bounds sharpen…
We define a novel class of distances between statistical multivariate distributions by modeling an optimal transport problem on their marginals with respect to a ground distance defined on their conditionals. These new distances are metrics…
We propose a variational approach to approximate measures with measures uniformly distributed over a 1 dimentional set. The problem consists in minimizing a Wasserstein distance as a data term with a regularization given by the length of…
Two geometrical structures have been extensively studied for a manifold of probability distributions. One is based on the Fisher information metric, which is invariant under reversible transformations of random variables, while the other is…
Given a determinate (multivariate) probability measure $\mu$, we characterize Gaussian mixtures $\nu\_\phi$ which minimize the Wasserstein distance $W\_2(\mu,\nu\_\phi)$ to $\mu$ when the mixing probability measure $\phi$ on the parameters…