English
Related papers

Related papers: Dimensionality Reduction has Quantifiable Imperfec…

200 papers

Sufficient dimension reduction (SDR) is continuing an active research field nowadays for high dimensional data. It aims to estimate the central subspace (CS) without making distributional assumption. To overcome the large-$p$-small-$n$…

Methodology · Statistics 2017-03-22 Hung Hung , Su-Yun Huang

We prove optimal sampling bounds achieving $(1\pm\varepsilon)$-relative error for a broad class of Lipschitz continuous classification loss functions under various regularization terms. This includes important functions such as logistic and…

Machine Learning · Computer Science 2026-05-25 Meysam Alishahi , Alexander Munteanu , Simon Omlor , Jeff M. Phillips

Topological Data Analysis methods can be useful for classification and clustering tasks in many different fields as they can provide two dimensional persistence diagrams that summarize important information about the shape of potentially…

Quantum Physics · Physics 2024-09-02 Bernardo Ameneyro , Rebekah Herrman , George Siopsis , Vasileios Maroulas

Assume that we observe i.i.d.~points lying close to some unknown $d$-dimensional $\mathcal{C}^k$ submanifold $M$ in a possibly high-dimensional space. We study the problem of reconstructing the probability distribution generating the…

Statistics Theory · Mathematics 2022-02-15 Vincent Divol

The Wasserstein metric has become increasingly important in many machine learning applications such as generative modeling, image retrieval and domain adaptation. Despite its appeal, it is often too costly to compute. This has motivated…

Machine Learning · Computer Science 2025-06-04 Jonathan Bobrutsky , Amit Moscovich

The Wasserstein metric or earth mover's distance (EMD) is a useful tool in statistics, machine learning and computer science with many applications to biological or medical imaging, among others. Especially in the light of increasingly…

Optimization and Control · Mathematics 2018-01-26 Jörn Schrieber , Dominic Schuhmacher , Carsten Gottschlich

Nonlinear dimensionality reduction methods are a popular tool for data scientists and researchers to visualize complex, high dimensional data. However, while these methods continue to improve and grow in number, it is often difficult to…

Machine Learning · Statistics 2019-09-04 Jonathan Johannemann , Robert Tibshirani

We study the structure of the support of a doubling measure by analyzing its self-similarity properties, which we estimate using a variant of the $L^1$ Wasserstein distance. We show that measure satisfying certain self-similarity conditions…

Metric Geometry · Mathematics 2014-11-11 Jonas Azzam , Guy David , Tatiana Toro

In this paper, we study the problem of sampling from a distribution under the constraint of differential privacy (DP). Prior works measure the utility of DP sampling with density ratio-based measures such as KL divergence. However, such…

Machine Learning · Statistics 2026-05-12 Shokichi Takakura , Seng Pei Liew , Satoshi Hasegawa

Distributionally robust optimization (DRO) has become a powerful framework for estimation under uncertainty, offering strong out-of-sample performance and principled regularization. In this paper, we propose a DRO-based method for linear…

Machine Learning · Statistics 2025-05-06 Liviu Aolaritei , Soroosh Shafiee , Florian Dörfler

The Wasserstein distance has become increasingly important in machine learning and deep learning. Despite its popularity, the Wasserstein distance is hard to approximate because of the curse of dimensionality. A recently proposed approach…

Machine Learning · Computer Science 2021-09-29 Minhui Huang , Shiqian Ma , Lifeng Lai

Optimal transport has gained much attention in image processing field, such as computer vision, image interpolation and medical image registration. Recently, Bredies et al. (ESAIM:M2AN 54:2351-2382, 2020) and Schmitzer et al. (IEEE T MED…

Numerical Analysis · Mathematics 2023-08-21 Yiming Gao

Making sense of Wasserstein distances between discrete measures in high-dimensional settings remains a challenge. Recent work has advocated a two-step approach to improve robustness and facilitate the computation of optimal transport, using…

Machine Learning · Computer Science 2019-09-04 François-Pierre Paty , Marco Cuturi

Deep neural networks (DNNs) exhibit an exceptional capacity for generalization in practical applications. This work aims to capture the effect and benefits of depth for supervised learning via information-theoretic generalization bounds. We…

Machine Learning · Computer Science 2025-05-09 Haiyun He , Ziv Goldfeld

Dimension reduction (DR) algorithms have proven to be extremely useful for gaining insight into large-scale high-dimensional datasets, particularly finding clusters in transcriptomic data. The initial phase of these DR methods often…

Machine Learning · Computer Science 2025-10-15 Yingfan Wang , Yiyang Sun , Haiyang Huang , Cynthia Rudin

In this study, we propose a high-performance disparity (depth) estimation method using dual-pixel (DP) images with few parameters. Conventional end-to-end deep-learning methods have many parameters but do not fully exploit disparity…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Teppei Kurita , Yuhi Kondo , Legong Sun , Takayuki Sasaki , Sho Nitta , Yasuhiro Hashimoto , Yoshinori Muramatsu , Yusuke Moriuchi

It was recently shown that under smoothness conditions, the squared Wasserstein distance between two distributions could be efficiently computed with appealing statistical error upper bounds. However, rather than the distance itself, the…

Machine Learning · Statistics 2021-12-30 Boris Muzellec , Adrien Vacher , Francis Bach , François-Xavier Vialard , Alessandro Rudi

\v{C}ech Persistence diagrams (PDs) are topological descriptors routinely used to capture the geometry of complex datasets. They are commonly compared using the Wasserstein distances $OT_{p}$; however, the extent to which PDs are stable…

Computational Geometry · Computer Science 2024-07-15 Charles Arnal , David Cohen-Steiner , Vincent Divol

Real-world data typically contain repeated and periodic patterns. This suggests that they can be effectively represented and compressed using only a few coefficients of an appropriate basis (e.g., Fourier, Wavelets, etc.). However, distance…

Machine Learning · Statistics 2014-05-26 Michail Vlachos , Nikolaos Freris , Anastasios Kyrillidis

Distributionally robust optimization (DRO) is an effective approach for data-driven decision-making in the presence of uncertainty. Geometric uncertainty due to sampling or localized perturbations of data points is captured by Wasserstein…

Machine Learning · Statistics 2023-11-10 Sloan Nietert , Ziv Goldfeld , Soroosh Shafiee