English
Related papers

Related papers: Wasserstein-based Kernel Principal Component Analy…

200 papers

Understanding proper distance measures between distributions is at the core of several learning tasks such as generative models, domain adaptation, clustering, etc. In this work, we focus on mixture distributions that arise naturally in…

Machine Learning · Computer Science 2019-10-30 Yogesh Balaji , Rama Chellappa , Soheil Feizi

Many applications in machine learning involve data represented as probability distributions. The emergence of such data requires radically novel techniques to design tractable gradient flows on probability distributions over this type of…

Machine Learning · Computer Science 2025-06-10 Clément Bonet , Christophe Vauthier , Anna Korba

The quantum Wasserstein distance (W-distance) is a fundamental metric for quantifying the distinguishability of quantum operations, with critical applications in quantum error correction. However, computing the W-distance remains…

Quantum Physics · Physics 2025-11-18 Changchun Feng , Xinyu Qiu , Laifa Tao , Lin Chen

Optimal transport has been very successful for various machine learning tasks; however, it is known to suffer from the curse of dimensionality. Hence, dimensionality reduction is desirable when applied to high-dimensional data with…

Machine Learning · Statistics 2025-07-21 Jie Wang , March Boedihardjo , Yao Xie

Comparing probability distributions is at the crux of many machine learning algorithms. Maximum Mean Discrepancies (MMD) and Wasserstein distances are two classes of distances between probability distributions that have attracted abundant…

Machine Learning · Statistics 2023-06-01 Titouan Vayer , Rémi Gribonval

The Wasserstein distance has emerged as a key metric to quantify distances between probability distributions, with applications in various fields, including machine learning, control theory, decision theory, and biological systems.…

Machine Learning · Computer Science 2026-02-10 Eduardo Figueiredo , Steven Adams , Luca Laurenti

In a variety of research areas, the weighted bag of vectors and the histogram are widely used descriptors for complex objects. Both can be expressed as discrete distributions. D2-clustering pursues the minimum total within-cluster variation…

Computation · Statistics 2017-01-10 Jianbo Ye , Panruo Wu , James Z. Wang , Jia Li

Using statistical learning methods to analyze stochastic simulation outputs can significantly enhance decision-making by uncovering relationships between different simulated systems and between a system's inputs and outputs. We focus on…

Methodology · Statistics 2026-05-28 Mohammadmahdi Ghasemloo , David J. Eckman

In this work we study systems consisting of a group of moving particles. In such systems, often some important parameters are unknown and have to be estimated from observed data. Such parameter estimation problems can often be solved via a…

Applications · Statistics 2023-07-11 Chen Cheng , Linjie Wen , Jinglai Li

Distances between probability distributions are a key component of many statistical machine learning tasks, from two-sample testing to generative modeling, among others. We introduce a novel distance between measures that compares them…

Machine Learning · Statistics 2025-07-09 Arturo Castellanos , Anna Korba , Pavlo Mozharovskyi , Hicham Janati

We introduce the observable Wasserstein distance, a framework for deriving lower bounds on the Wasserstein distance between probability measures on Polish metric spaces, designed to bypass the computational intractability of exact optimal…

Metric Geometry · Mathematics 2026-05-12 Edivaldo Lopes dos Santos , Leandro Vicente Mauri , Washington Mio , Tom Needham

This paper uses sample data to study the problem of comparing populations on finite-dimensional parallelizable Riemannian manifolds and more general trivial vector bundles. Utilizing triviality, our framework represents populations as…

Methodology · Statistics 2023-11-29 Michael Wilson , Tom Needham , Chiwoo Park , Suprateek Kundu , Anuj Srivastava

Model-based clustering is widely-used in a variety of application areas. However, fundamental concerns remain about robustness. In particular, results can be sensitive to the choice of kernel representing the within-cluster data density.…

Machine Learning · Statistics 2019-06-27 Leo L Duan , David B Dunson

In a high-dimensional regression framework, we study consequences of the naive two-step procedure where first the dimension of the input variables is reduced and second, the reduced input variables are used to predict the output variable…

Machine Learning · Statistics 2023-11-28 Stephan Eckstein , Armin Iske , Mathias Trabs

This paper deals with a clustering algorithm for histogram data based on a Self-Organizing Map (SOM) learning. It combines a dimension reduction by SOM and the clustering of the data in a reduced space. Related to the kind of data, a…

Machine Learning · Computer Science 2021-09-10 Guénaël Cabanes , Younès Bennani , Rosanna Verde , Antonio Irpino

Learning conditional densities and identifying factors that influence the entire distribution are vital tasks in data-driven applications. Conventional approaches work mostly with summary statistics, and are hence inadequate for a…

Methodology · Statistics 2022-09-13 Chengliang Tang , Nathan Lenssen , Ying Wei , Tian Zheng

The topological patterns exhibited by many real-world networks motivate the development of topology-based methods for assessing the similarity of networks. However, extracting topological structure is difficult, especially for large and…

Machine Learning · Computer Science 2022-03-15 Tananun Songdechakraiwut , Bryan M. Krause , Matthew I. Banks , Kirill V. Nourski , Barry D. Van Veen

Persistence diagrams (PDs) play a key role in topological data analysis (TDA), in which they are routinely used to describe topological properties of complicated shapes. PDs enjoy strong stability properties and have proven their utility in…

Computational Geometry · Computer Science 2017-11-10 Mathieu Carrière , Marco Cuturi , Steve Oudot

Change point detection for time series analysis is a difficult and important problem in applied statistics, for which a variety of approaches have been developed in the past several decades. Here, the Wasserstein metric is employed as a…

Statistics Theory · Mathematics 2026-03-03 David Gentile , Joshua Huang , James M. Murphy

With the rapid development of online payment platforms, it is now possible to record massive transaction data. Clustering on transaction data significantly contributes to analyzing merchants' behavior patterns. This enables payment…

Methodology · Statistics 2023-02-17 Yingqiu Zhu , Danyang Huang , Bo Zhang