中文
相关论文

相关论文: Minimum intrinsic dimension scaling for entropic o…

200 篇论文

Many algorithms in machine learning and computational geometry require, as input, the intrinsic dimension of the manifold that supports the probability distribution of the data. This parameter is rarely known and therefore has to be…

统计理论 · 数学 2020-01-01 Jisu Kim , Alessandro Rinaldo , Larry Wasserman

The notion of entropy-regularized optimal transport, also known as Sinkhorn divergence, has recently gained popularity in machine learning and statistics, as it makes feasible the use of smoothed optimal transportation distances for data…

统计理论 · 数学 2019-11-05 Jérémie Bigot , Elsa Cazelles , Nicolas Papadakis

We derive nearly tight and non-asymptotic convergence bounds for solutions of entropic semi-discrete optimal transport. These bounds quantify the stability of the dual solutions of the regularized problem (sometimes called Sinkhorn…

人工智能 · 计算机科学 2022-05-05 Alex Delalande

Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a…

机器学习 · 统计学 2018-03-20 Elena Facco , Maria d'Errico , Alex Rodriguez , Alessandro Laio

Optimal transport (OT) serves as a natural framework for comparing probability measures, with applications in statistics, machine learning, and applied mathematics. Alas, statistical estimation and exact computation of the OT distances…

统计理论 · 数学 2024-05-14 Tao Wang , Ziv Goldfeld

One of the founding paradigms of machine learning is that a small number of variables is often sufficient to describe high-dimensional data. The minimum number of variables required is called the intrinsic dimension (ID) of the data.…

机器学习 · 统计学 2020-07-14 Michele Allegra , Elena Facco , Francesco Denti , Alessandro Laio , Antonietta Mira

Optimal transport has emerged as a fundamental methodology with applications spanning multiple research areas in recent years. However, the convergence rate of the empirical estimator to its population counterpart suffers from the curse of…

统计理论 · 数学 2025-10-06 Jiaping Yang , Yunxin Zhang

The real-life data have a complex and non-linear structure due to their nature. These non-linearities and the large number of features can usually cause problems such as the empty-space phenomenon and the well-known curse of dimensionality.…

机器学习 · 计算机科学 2025-03-13 Kadir Özçoban , Murat Manguoğlu , Emrullah Fatih Yetkin

We study the optimal transport problem for $d>2$ discrete measures. This is a linear programming problem on $d$-tensors. It gives a way to compute a "distance" between two sets of discrete measures. We introduce an entropic regularization…

计算机视觉与模式识别 · 计算机科学 2021-07-27 Shmuel Friedland

The intrinsic dimensionality refers to the ``true'' dimensionality of the data, as opposed to the dimensionality of the data representation. For example, when attributes are highly correlated, the intrinsic dimensionality can be much lower…

机器学习 · 统计学 2020-11-30 Erik Thordsen , Erich Schubert

We propose a new method for estimating the intrinsic dimension of a dataset by applying the principle of regularized maximum likelihood to the distances between close neighbors. We propose a regularization scheme which is motivated by…

机器学习 · 计算机科学 2012-03-19 Mithun Das Gupta , Thomas S. Huang

We investigate the small regularization limit of entropic optimal transport when the cost function is the Euclidean distance in dimensions $d > 1$, and the marginal measures are absolutely continuous with respect to the Lebesgue measure.…

概率论 · 数学 2025-08-15 Shrey Aryan , Promit Ghosal

Multidimensional scaling (MDS) is the act of embedding proximity information about a set of $n$ objects in $d$-dimensional Euclidean space. As originally conceived by the psychometric community, MDS was concerned with embedding a fixed set…

机器学习 · 统计学 2024-12-12 Michael W. Trosset , Carey E. Priebe

Entropically regularized optimal transport between probability measures supported on compact subsets of Euclidean space admits a representation as an information projection under moment inequality constraints. Exploiting this structure, I…

统计理论 · 数学 2026-01-15 Rami V. Tabri

We study the regularity properties of the minimisers of entropic optimal transport providing a natural analogue of the $\varepsilon$-regularity theory of quadratic optimal transport in the entropic setting. More precisely, we show that if…

偏微分方程分析 · 数学 2025-01-14 Rishabh S. Gvalani , Lukas Koch

Most entropy measures depend on the spread of the probability distribution over the sample space $\mathcal{X}$, and the maximum entropy achievable scales proportionately with the sample space cardinality $|\mathcal{X}|$. For a finite…

机器学习 · 计算机科学 2023-05-25 Rohan Ghosh , Mehul Motani

The Intrinsic Dimension (ID) is a key concept in unsupervised learning and feature selection, as it is a lower bound to the number of variables which are necessary to describe a system. However, in almost any real-world dataset the ID…

机器学习 · 统计学 2026-04-02 Antonio Di Noia , Iuri Macocco , Aldo Glielmo , Alessandro Laio , Antonietta Mira

Let $\mathcal{M}$ be a smooth submanifold of $\mathbb{R}^n$ equipped with the Euclidean (chordal) metric. This note considers the smallest dimension $m$ for which there exists a bi-Lipschitz function $f: \mathcal{M} \mapsto \mathbb{R}^m$…

数值分析 · 数学 2021-05-31 Mark Iwen , Arman Tavakoli , Benjamin Schmidt

Quantum many-body states that frequently appear in physics often obey an entropy scaling law, meaning that an entanglement entropy of a subsystem can be expressed as a sum of terms that scale linearly with its volume and area, plus a…

量子物理 · 物理学 2021-05-26 Isaac H. Kim

High-dimensional data are ubiquitous in contemporary science and finding methods to compress them is one of the primary goals of machine learning. Given a dataset lying in a high-dimensional space (in principle hundreds to several thousands…

机器学习 · 计算机科学 2020-03-24 Vittorio Erba , Marco Gherardi , Pietro Rotondo
‹ 上一页 1 2 3 10 下一页 ›