English
Related papers

Related papers: Manifold Hypothesis in Data Analysis: Double Geome…

200 papers

Manifolds discovered by machine learning models provide a compact representation of the underlying data. Geodesics on these manifolds define locally length-minimising curves and provide a notion of distance, which are key for reduced-order…

Machine Learning · Computer Science 2023-05-25 Daniel Kelshaw , Luca Magri

The general aim of manifold estimation is reconstructing, by statistical methods, an $m$-dimensional compact manifold $S$ on ${\mathbb R}^d$ (with $m\leq d$) or estimating some relevant quantities related to the geometric properties of $S$.…

Statistics Theory · Mathematics 2014-11-13 José R. Berrendero , Alejandro Cholaquidis , Antonio Cuevas , Ricardo Fraiman

Manifold learning (ML), known also as non-linear dimension reduction, is a set of methods to find the low dimensional structure of data. Dimension reduction for large, high dimensional data is not merely a way to reduce the data; the new…

Machine Learning · Statistics 2023-11-08 Marina Meilă , Hanyu Zhang

In this work, we investigate Riemannian geometry based dimensionality reduction methods that respect the underlying manifold structure of the data. In particular, we focus on Principal Geodesic Analysis (PGA) as a nonlinear generalization…

Machine Learning · Computer Science 2026-02-06 Alaa El Ichi , Khalide Jbilou

The manifold hypothesis posits that high-dimensional data typically resides on low-dimensional sub spaces. In this paper, we assume manifold hypothesis to investigate graph-based semi-supervised learning methods. In particular, we examine…

Machine Learning · Computer Science 2025-11-18 Mary Chriselda Antony Oliver , Michael Roberts , Carola-Bibiane Schönlieb , Matthew Thorpe

A common observation in data-driven applications is that high-dimensional data have a low intrinsic dimension, at least locally. In this work, we consider the problem of point estimation for manifold-valued data. Namely, given a finite set…

Statistics Theory · Mathematics 2025-03-11 Yariv Aizenbud , Barak Sober

Two-sample tests are important areas aiming to determine whether two collections of observations follow the same distribution or not. We propose two-sample tests based on integral probability metric (IPM) for high-dimensional samples…

Machine Learning · Statistics 2023-04-21 Jie Wang , Minshuo Chen , Tuo Zhao , Wenjing Liao , Yao Xie

There are many distance-based methods for classification and clustering, and for data with a high number of dimensions and a lower number of observations, processing distances is computationally advantageous compared to the raw data matrix.…

Methodology · Statistics 2020-06-25 Christian Hennig

For manifold learning, it is assumed that high-dimensional sample/data points are embedded on a low-dimensional manifold. Usually, distances among samples are computed to capture an underlying data structure. Here we propose a metric…

Machine Learning · Computer Science 2019-09-20 Fenglei Fan , Ziyu Su , Yueyang Teng , Ge Wang

We propose a new method for estimating the intrinsic dimension of a dataset by applying the principle of regularized maximum likelihood to the distances between close neighbors. We propose a regularization scheme which is motivated by…

Machine Learning · Computer Science 2012-03-19 Mithun Das Gupta , Thomas S. Huang

Laplacian-based methods are popular for the dimensionality reduction of data lying in $\mathbb{R}^N$. Several theoretical results for these algorithms depend on the fact that the Euclidean distance locally approximates the geodesic distance…

Machine Learning · Computer Science 2025-09-24 Liane Xu , Amit Singer

It is a standard assumption that datasets in high dimension have an internal structure which means that they in fact lie on, or near, subsets of a lower dimension. In many instances it is important to understand the real dimension of the…

Machine Learning · Statistics 2025-07-21 James A. D. Binnie , Paweł Dłotko , John Harvey , Jakub Malinowski , Ka Man Yim

Scientists and engineers rely on accurate mathematical models to quantify the objects of their studies, which are often high-dimensional. Unfortunately, high-dimensional models are inherently difficult, i.e. when observations are sparse or…

Machine Learning · Computer Science 2018-02-13 Robert A. Bridges , Chris Felder , Chelsey Hoff

There has been a growing interest in statistical inference from data satisfying the so-called manifold hypothesis, assuming data points in the high-dimensional ambient space to lie in close vicinity of a submanifold of much lower dimension.…

Methodology · Statistics 2023-01-04 Rong Tang , Yun Yang

Statistical analysis of a graph often starts with embedding, the process of representing its nodes as points in space. How to choose the embedding dimension is a nuanced decision in practice, but in theory a notion of true dimension is…

Machine Learning · Statistics 2021-01-06 Patrick Rubin-Delanchy

The real-life data have a complex and non-linear structure due to their nature. These non-linearities and the large number of features can usually cause problems such as the empty-space phenomenon and the well-known curse of dimensionality.…

Machine Learning · Computer Science 2025-03-13 Kadir Özçoban , Murat Manguoğlu , Emrullah Fatih Yetkin

Detecting the dimension of a hidden manifold from a point sample has become an important problem in the current data-driven era. Indeed, estimating the shape dimension is often the first step in studying the processes or phenomena…

Computational Geometry · Computer Science 2014-05-15 Tamal K. Dey , Fengtao Fan , Yusu Wang

Function approximation based on data drawn randomly from an unknown distribution is an important problem in machine learning. The manifold hypothesis assumes that the data is sampled from an unknown submanifold of a high dimensional…

Machine Learning · Computer Science 2024-08-20 H. N. Mhaskar , Ryan O'Dowd

Experimental sciences have come to depend heavily on our ability to organize and interpret high-dimensional datasets. Natural laws, conservation principles, and inter-dependencies among observed variables yield geometric structure, with…

Quantum Physics · Physics 2022-12-15 Akshat Kumar , Mohan Sarovar

High dimensional data can have a surprising property: pairs of data points may be easily separated from each other, or even from arbitrary subsets, with high probability using just simple linear classifiers. However, this is more of a rule…

Machine Learning · Computer Science 2023-11-15 Oliver J. Sutton , Qinghua Zhou , Alexander N. Gorban , Ivan Y. Tyukin