Related papers: A new metric between distributions of point proces…
This study addresses a class of linear mixed-integer programming (MILP) problems that involve uncertainty in the objective function parameters. The parameters are assumed to form a random vector, whose probability distribution can only be…
We establish a new global endpoint Sobolev inequality for measures that extends the classical theorem of Meyers-Ziemer by placing a maximal function on the right-hand side. This result has several significant consequences. It extends…
Approximating a probability distribution using a set of particles is a fundamental problem in machine learning and statistics, with applications including clustering and quantization. Formally, we seek a weighted mixture of Dirac measures…
1. Complex systems of moving and interacting objects are ubiquitous in the natural and social sciences. Predicting their behavior often requires models that mimic these systems with sufficient accuracy, while accounting for their inherent…
A common feature of methods for analyzing samples of probability density functions is that they respect the geometry inherent to the space of densities. Once a metric is specified for this space, the Fr\'echet mean is typically used to…
The proliferation of large data sets and Bayesian inference techniques motivates demand for better data sparsification. Coresets provide a principled way of summarizing a large dataset via a smaller one that is guaranteed to match the…
This paper studies distributional model risk in marginal problems, where each marginal measure is assumed to lie in a Wasserstein ball centered at a fixed reference measure with a given radius. Theoretically, we establish several…
Let $M$ be a $d$-dimensional connected compact Riemannian manifold with boundary $\partial M$, let $V\in C^2(M)$ such that $\mu({\rm d} x):={\rm e}^{V(x)}{\rm d} x$ is a probability measure, and let $X_t$ be the diffusion process generated…
We construct a system of interacting two-sided Bessel processes on the unit interval and show that the associated empirical measure process converges to the Wasserstein Diffusion, assuming that Markov uniqueness holds for the generating…
We consider finite point subsets (distributions) in compact metric spaces. Non-trivial bounds for sums of distances between points of distributions and for discrepancies of distributions in metric balls are given in the case of general…
Increasingly complex data analysis tasks motivate the study of the dependency of distributions of multivariate continuous random variables on scalar or vector predictors. Statistical regression models for distributional responses so far…
This paper studies the sensitivity analysis of mass-action systems against their diffusion approximations, particularly the dependence on population sizes. As a continuous time Markov chain, a mass-action system can be described by a…
In this work, we provide non-asymptotic bounds for the average speed of convergence of the empirical measure in the law of large numbers, in Wasserstein distance. We also consider occupation measures of ergodic Markov chains. One motivation…
Wasserstein barycenters define averages of probability measures in a geometrically meaningful way. Their use is increasingly popular in applied fields, such as image, geometry or language processing. In these fields however, the probability…
The object of study in this paper is the expected $2$-Wasserstein distance between the empirical measures of several point processes and their respective limit. For this, the main tool developed is a smoothing procedure in Euclidean spaces…
We propose a novel family of test statistics to detect the presence of changepoints in a sequence of dependent, possibly multivariate, functional-valued observations. Our approach allows to test for a very general class of changepoints,…
A common way to discretize a probability measure is to use an empirical measure as a discrete approximation. But how far from being optimal is this approximation in the p-Wasserstein distance? In this paper, we study this question in two…
Regression analysis with probability measures as input predictors and output response has recently drawn great attention. However, it is challenging to handle multiple input probability measures due to the non-flat Riemannian geometry of…
The post-processing approaches are becoming prominent techniques to enhance machine learning models' fairness because of their intuitiveness, low computational cost, and excellent scalability. However, most existing post-processing methods…
Problem definition: A key challenge in supervised learning is data scarcity, which can cause prediction models to overfit to the training data and perform poorly out of sample. A contemporary approach to combat overfitting is offered by…