English
Related papers

Related papers: Robust clustering tools based on optimal transport…

200 papers

This paper is concerned by the study of barycenters for random probability measures in the Wasserstein space. Using a duality argument, we give a precise characterization of the population barycenter for various parametric classes of random…

Statistics Theory · Mathematics 2017-11-30 Jérémie Bigot , Thierry Klein

The Wasserstein barycenter extends the Euclidean mean to the space of probability measures by minimizing the weighted sum of squared 2-Wasserstein distances. We develop a free-support algorithm for computing Wasserstein barycenters that…

Machine Learning · Statistics 2025-09-17 Kisung You

The increasing availability of granular and big data on various objects of interest has made it necessary to develop methods for condensing this information into a representative and intelligible map. Financial regulation is a field that…

Machine Learning · Statistics 2025-07-08 Lorenz Riess , Mathias Beiglböck , Johannes Temme , Andreas Wolf , Julio Backhoff

Wasserstein Barycenter is a principled approach to represent the weighted mean of a given set of probability distributions, utilizing the geometry induced by optimal transport. In this work, we present a novel scalable algorithm to…

Machine Learning · Computer Science 2021-11-30 Jiaojiao Fan , Amirhossein Taghvaei , Yongxin Chen

Wasserstein barycenters correspond to optimal solutions of transportation problems for several marginals, and as such have a wide range of applications ranging from economics to statistics and computer science. When the marginal probability…

Optimization and Control · Mathematics 2015-08-11 Ethan Anderes , Steffen Borgwardt , Jacob Miller

The Wasserstein barycenter problem is to compute the average of $m$ given probability measures, which has been widely studied in many different areas; however, real-world data sets are often noisy and huge, which impedes its applications in…

Machine Learning · Computer Science 2023-12-27 Xu Wang , Jiawei Huang , Qingyuan Yang , Jinpeng Zhang

The fairness of clustering algorithms has gained widespread attention across various areas, including machine learning, In this paper, we study fair $k$-means clustering in Euclidean space. Given a dataset comprising several groups, the…

Machine Learning · Computer Science 2024-12-10 Shihong Song , Guanlin Mo , Qingyuan Yang , Hu Ding

We present a novel method for efficiently computing optimal transport maps and Wasserstein barycenters in high-dimensional spaces. Our approach uses conditional normalizing flows to approximate the input distributions as invertible…

Machine Learning · Statistics 2025-05-29 Gabriele Visentin , Patrick Cheridito

The Bayesian approach to clustering is often appreciated for its ability to provide uncertainty in the partition structure. However, summarizing the posterior distribution over the clustering structure can be challenging, due the discrete,…

Computation · Statistics 2026-01-26 Cecilia Balocchi , Sara Wade

We give an efficient algorithm for robustly clustering of a mixture of two arbitrary Gaussians, a central open problem in the theory of computationally efficient robust estimation, assuming only that the the means of the component Gaussians…

Data Structures and Algorithms · Computer Science 2020-06-02 He Jia , Santosh Vempala

Flexible Bayesian models are typically constructed using limits of large parametric models with a multitude of parameters that are often uninterpretable. In this article, we offer a novel alternative by constructing an exponentially tilted…

Methodology · Statistics 2023-03-20 Abhisek Chakraborty , Anirban Bhattacharya , Debdeep Pati

Wasserstein barycenters provide a geometric notion of the weighted average of probability measures based on optimal transport. In this paper, we present a scalable algorithm to compute Wasserstein-2 barycenters given sample access to the…

Machine Learning · Computer Science 2022-01-02 Alexander Korotin , Lingxiao Li , Justin Solomon , Evgeny Burnaev

The problem of rapid and automated detection of distinct market regimes is a topic of great interest to financial mathematicians and practitioners alike. In this paper, we outline an unsupervised learning algorithm for clustering financial…

Computational Finance · Quantitative Finance 2021-10-25 Blanka Horvath , Zacharia Issa , Aitor Muguruza

We address general-shaped clustering problems under very weak parametric assumptions with a two-step hybrid robust clustering algorithm based on trimmed k-means and hierarchical agglomeration. The algorithm has low computational complexity…

Methodology · Statistics 2022-01-19 Luca Insolia , Domenico Perrotta

Consider a multi-agent system whereby each agent has an initial probability measure. In this paper, we propose a distributed algorithm based upon stochastic, asynchronous and pairwise exchange of information and displacement interpolation…

Systems and Control · Electrical Eng. & Systems 2022-02-28 Pedro Cisneros-Velarde , Francesco Bullo

One key use of k-means clustering is to identify cluster prototypes which can serve as representative points for a dataset. However, a drawback of using k-means cluster centers as representative points is that such points distort the…

Machine Learning · Statistics 2019-11-15 Arvind Krishna , Simon Mak , Roshan Joseph

The primary choice to summarize a finite collection of random objects is by using measures of central tendency, such as mean and median. In the field of optimal transport, the Wasserstein barycenter corresponds to the Fr\'{e}chet or…

Methodology · Statistics 2025-09-03 Kisung You , Dennis Shung , Mauro Giuffrè

This paper studies the statistical estimation of exact Wasserstein barycenters. Existing non-asymptotic results for empirical barycenters exhibit a severe curse of dimensionality. Motivated by the semi-dual formulation of the barycenter…

Statistics Theory · Mathematics 2026-05-06 Pengtao Li , Changbo Zhu , Xiaohui Chen

We present a framework to simultaneously align and smooth data in the form of multiple point clouds sampled from unknown densities with support in a d-dimensional Euclidean space. This work is motivated by applications in bioinformatics…

Methodology · Statistics 2019-08-28 Jérémie Bigot , Elsa Cazelles , Nicolas Papadakis

Clustering is a fundamental tool in unsupervised learning, used to group objects by distinguishing between similar and dissimilar features of a given data set. One of the most common clustering algorithms is k-means. Unfortunately, when…

Machine Learning · Statistics 2021-08-17 Olga Dorabiala , J. Nathan Kutz , Aleksandr Aravkin