English
Related papers

Related papers: Minimum intrinsic dimension scaling for entropic o…

200 papers

The ability to represent and compare machine learning models is crucial in order to quantify subtle model changes, evaluate generative models, and gather insights on neural network architectures. Existing techniques for comparing data…

Comparing probability distributions is a fundamental problem in data sciences. Simple norms and divergences such as the total variation and the relative entropy only compare densities in a point-wise manner and fail to capture the geometric…

We define a novel class of distances between statistical multivariate distributions by modeling an optimal transport problem on their marginals with respect to a ground distance defined on their conditionals. These new distances are metrics…

Machine Learning · Computer Science 2020-11-03 Frank Nielsen , Ke Sun

High dimensional data can have a surprising property: pairs of data points may be easily separated from each other, or even from arbitrary subsets, with high probability using just simple linear classifiers. However, this is more of a rule…

Machine Learning · Computer Science 2023-11-15 Oliver J. Sutton , Qinghua Zhou , Alexander N. Gorban , Ivan Y. Tyukin

Dimension reduction algorithms are a crucial part of many data science pipelines, including data exploration, feature creation and selection, and denoising. Despite their wide utilization, many non-linear dimension reduction algorithms are…

Machine Learning · Statistics 2024-08-06 Ryan Murray , Adam Pickarski

We study the existing algorithms that solve the multidimensional martingale optimal transport. Then we provide a new algorithm based on entropic regularization and Newton's method. Then we provide theoretical convergence rate results and we…

Probability · Mathematics 2018-12-31 Hadrien De March

Manifold learning is a central task in modern statistics and data science. Many datasets (cells, documents, images, molecules) can be represented as point clouds embedded in a high dimensional ambient space, however the degrees of freedom…

Machine Learning · Statistics 2025-02-18 Stephen Zhang , Gilles Mordant , Tetsuya Matsumoto , Geoffrey Schiebinger

Deep learning has had tremendous success at learning low-dimensional representations of high-dimensional data. This success would be impossible if there was no hidden low-dimensional structure in data of interest; this existence is posited…

Most of the existing methods for estimating the local intrinsic dimension of a data distribution do not scale well to high-dimensional data. Many of them rely on a non-parametric nearest neighbors approach which suffers from the curse of…

Rotation Averaging is a non-convex optimization problem that determines orientations of a collection of cameras from their images of a 3D scene. The problem has been studied using a variety of distances and robustifiers. The intrinsic (or…

Computer Vision and Pattern Recognition · Computer Science 2020-03-19 Kyle Wilson , David Bindel

The empirical optimal transport (OT) cost between two probability measures from random data is a fundamental quantity in transport based data analysis. In this work, we derive novel guarantees for its convergence rate when the involved…

Statistics Theory · Mathematics 2022-02-22 Shayan Hundrieser , Thomas Staudt , Axel Munk

We prove several fundamental statistical bounds for entropic OT with the squared Euclidean cost between subgaussian probability measures in arbitrary dimension. First, through a new sample complexity result we establish the rate of…

Statistics Theory · Mathematics 2019-05-31 Gonzalo Mena , Jonathan Weed

We prove optimal bounds for the convergence rate of ordinal embedding (also known as non-metric multidimensional scaling) in the 1-dimensional case. The examples witnessing optimality of our bounds arise from a result in additive number…

Statistics Theory · Mathematics 2019-05-01 Jordan S. Ellenberg , Lalit Jain

For a given metric measure space $(X,d,\mu)$ we consider finite samples of points, calculate the matrix of distances between them and then reconstruct the points in some finite-dimensional space using the multidimensional scaling (MDS)…

Metric Geometry · Mathematics 2022-08-02 Alexey Kroshnin , Eugene Stepanov , Dario Trevisan

This paper considers estimation and inference in semiparametric econometric models. Standard procedures estimate the model based on an independence restriction that induces a minimum distance between a joint cumulative distribution function…

Statistics Theory · Mathematics 2014-12-09 Zhengyuan Gao , Antonio Galvao

We propose a unified scaling theory of entanglement entropy in the confinements of finite bond dimensions, dynamics and system sizes. Within the theory, the finite-entanglement scaling introduced recently is generalized to the dynamics…

Statistical Mechanics · Physics 2018-12-26 Xuanmin Cao , Qijun Hu , Fan Zhong

High-dimensional datasets often exhibit low-dimensional geometric structures, as suggested by the manifold hypothesis, which implies that data lie on a smooth manifold embedded in a higher-dimensional ambient space. While this insight…

Machine Learning · Computer Science 2025-07-11 Paola Causin , Alessio Marta

Optimal transport induces the Earth Mover's (Wasserstein) distance between probability distributions, a geometric divergence that is relevant to a wide range of problems. Over the last decade, two relaxations of optimal transport have been…

Optimization and Control · Mathematics 2023-01-18 Thibault Séjourné , Jean Feydy , François-Xavier Vialard , Alain Trouvé , Gabriel Peyré

Solving large scale Optimal Transport (OT) in machine learning typically relies on sampling measures to obtain a tractable discrete problem. While the discrete solver's accuracy is controllable, the rate of convergence of the discretization…

Machine Learning · Statistics 2026-02-05 Ferdinand Genans , Olivier Wintenberger

We study asymptotic lower and upper bounds for the sizes of constant dimension codes with respect to the subspace or injection distance, which is used in random linear network coding. In this context we review known upper bounds and show…

Combinatorics · Mathematics 2017-12-06 Daniel Heinlein , Sascha Kurz