中文
相关论文

相关论文: An axiomatic approach to intrinsic dimension of a …

200 篇论文

We propose an axiomatic approach to the concept of an intrinsic dimension of a dataset, based on a viewpoint of geometry of high-dimensional structures. Our first axiom postulates that high values of dimension be indicative of the presence…

机器学习 · 计算机科学 2016-11-17 Vladimir Pestov

The curse of dimensionality is a phenomenon frequently observed in machine learning (ML) and knowledge discovery (KD). There is a large body of literature investigating its origin and impact, using methods from mathematics as well as from…

人工智能 · 计算机科学 2022-04-22 Tom Hanika , Friedrich Martin Schneider , Gerd Stumme

The concept of dimension is essential to grasp the complexity of data. A naive approach to determine the dimension of a dataset is based on the number of attributes. More sophisticated methods derive a notion of intrinsic dimension (ID)…

机器学习 · 计算机科学 2023-04-18 Maximilian Stubbemann , Tom Hanika , Friedrich Martin Schneider

The real-life data have a complex and non-linear structure due to their nature. These non-linearities and the large number of features can usually cause problems such as the empty-space phenomenon and the well-known curse of dimensionality.…

机器学习 · 计算机科学 2025-03-13 Kadir Özçoban , Murat Manguoğlu , Emrullah Fatih Yetkin

The curse of dimensionality in the realm of association rules is twofold. Firstly, we have the well known exponential increase in computational complexity with increasing item set size. Secondly, there is a \emph{related curse} concerned…

人工智能 · 计算机科学 2018-05-16 Tom Hanika , Friedrich Martin Schneider , Gerd Stumme

Real-world datasets are often of high dimension and effected by the curse of dimensionality. This hinders their comprehensibility and interpretability. To reduce the complexity feature selection aims to identify features that are crucial to…

机器学习 · 计算机科学 2023-04-18 Maximilian Stubbemann , Tobias Hille , Tom Hanika

High dimensional data can have a surprising property: pairs of data points may be easily separated from each other, or even from arbitrary subsets, with high probability using just simple linear classifiers. However, this is more of a rule…

机器学习 · 计算机科学 2023-11-15 Oliver J. Sutton , Qinghua Zhou , Alexander N. Gorban , Ivan Y. Tyukin

Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a…

机器学习 · 统计学 2018-03-20 Elena Facco , Maria d'Errico , Alex Rodriguez , Alessandro Laio

Real world-datasets characterized by discrete features are ubiquitous: from categorical surveys to clinical questionnaires, from unweighted networks to DNA sequences. Nevertheless, the most common unsupervised dimensional reduction methods…

机器学习 · 统计学 2023-03-14 Iuri Macocco , Aldo Glielmo , Jacopo Grilli , Alessandro Laio

The size of datasets has been increasing rapidly both in terms of number of variables and number of events. As a result, the empty space phenomenon and the curse of dimensionality complicate the extraction of useful information. But, in…

数据分析、统计与概率 · 物理学 2015-05-07 Jean Golay , Mikhail Kanevski

The manifold hypothesis suggests that high-dimensional data often lie on or near a low-dimensional manifold. Estimating the dimension of this manifold is essential for leveraging its structure, yet existing work on dimension estimation is…

机器学习 · 计算机科学 2026-04-02 Zelong Bi , Pierre Lafaye de Micheaux

It is widely believed that natural image data exhibits low-dimensional structure despite the high dimensionality of conventional pixel representations. This idea underlies a common intuition for the remarkable success of deep learning in…

计算机视觉与模式识别 · 计算机科学 2021-04-20 Phillip Pope , Chen Zhu , Ahmed Abdelkader , Micah Goldblum , Tom Goldstein

It has long been noticed that high dimension data exhibits strange patterns. This has been variously interpreted as either a "blessing" or a "curse", causing uncomfortable inconsistencies in the literature. We propose that these patterns…

计算机视觉与模式识别 · 计算机科学 2020-03-18 Wen-Yan Lin

The rapid growth of high-dimensional datasets across various scientific domains has created a pressing need for new statistical methods to compare distributions supported on their underlying structures. Assessing similarity between datasets…

统计理论 · 数学 2025-11-27 Hongrui Chen , Rong Ma

The discovering of low-dimensional manifolds in high-dimensional data is one of the main goals in manifold learning. We propose a new approach to identify the effective dimension (intrinsic dimension) of low-dimensional manifolds. The scale…

统计理论 · 数学 2008-03-17 Xiaohui Wang , J. S. Marron

To gain insight into the mechanisms behind machine learning methods, it is crucial to establish connections among the features describing data points. However, these correlations often exhibit a high-dimensional and strongly nonlinear…

机器学习 · 计算机科学 2025-03-04 Lorenzo Basile , Santiago Acevedo , Luca Bortolussi , Fabio Anselmi , Alex Rodriguez

Many approaches in the field of machine learning and data analysis rely on the assumption that the observed data lies on lower-dimensional manifolds. This assumption has been verified empirically for many real data sets. To make use of this…

机器学习 · 计算机科学 2022-09-27 Erik Thordsen , Erich Schubert

It is a standard assumption that datasets in high dimension have an internal structure which means that they in fact lie on, or near, subsets of a lower dimension. In many instances it is important to understand the real dimension of the…

机器学习 · 统计学 2025-07-21 James A. D. Binnie , Paweł Dłotko , John Harvey , Jakub Malinowski , Ka Man Yim

We consider the problem of reconstructing the intrinsic geometry of a manifold from noisy pairwise distance observations. Specifically, let $M$ denote a diameter 1 d-dimensional manifold and $\mu$ a probability measure on $M$ that is…

机器学习 · 统计学 2025-11-18 Charles Fefferman , Jonathan Marty , Kevin Ren

Difficulties in replication and reproducibility of empirical evidences in machine learning research have become a prominent topic in recent years. Ensuring that machine learning research results are sound and reliable requires…

机器学习 · 计算机科学 2024-03-20 Tobias Hille , Maximilian Stubbemann , Tom Hanika
‹ 上一页 1 2 3 10 下一页 ›