English
Related papers

Related papers: What is the $\textit{intrinsic}$ dimension of your…

200 papers

We perform a deeper analysis of an axiomatic approach to the concept of intrinsic dimension of a dataset proposed by us in the IJCNN'07 paper (arXiv:cs/0703125). The main features of our approach are that a high intrinsic dimension of a…

Information Retrieval · Computer Science 2009-11-17 Vladimir Pestov

Axis-aligned subspace clustering generally entails searching through enormous numbers of subspaces (feature combinations) and evaluation of cluster quality within each subspace. In this paper, we tackle the problem of identifying subsets of…

Machine Learning · Computer Science 2019-07-17 Ruben Becker , Imane Hafnaoui , Michael E. Houle , Pan Li , Arthur Zimek

How to extract useful insights from data is always a challenge, especially if the data is multidimensional. Often, the data can be organized according to certain hierarchical structure that are stemmed either from data collection process or…

Applications · Statistics 2016-04-21 Kun Yang , Wing Hung Wong

Intrinsic dimension and differential entropy estimators are studied in this paper, including their systematic bias. A pragmatic approach for joint estimation and bias correction of these two fundamental measures is proposed. Shared steps on…

Machine Learning · Statistics 2020-05-01 Jugurta Montalvão , Jânio Canuto , Luiz Miranda

Cognitive Dimensions is a framework for analyzing human-computer interaction. It is used for meta-analysis, that is, for talking about characteristics of systems without getting bogged down in details of a particular implementation. In this…

Human-Computer Interaction · Computer Science 2009-08-26 Gene Golovchinsky

Data visualization is the process by which data of any size or dimensionality is processed to produce an understandable set of data in a lower dimensionality, allowing it to be manipulated and understood more easily by people. The goal of…

Graphics · Computer Science 2021-07-06 Alexander Kiefer , Md. Khaledur Rahman

Although pretrained language models can be fine-tuned to produce state-of-the-art results for a very wide range of language understanding tasks, the dynamics of this process are not well understood, especially in the low data regime. Why…

Machine Learning · Computer Science 2020-12-25 Armen Aghajanyan , Luke Zettlemoyer , Sonal Gupta

With the advancement of data science, the collection of increasingly complex datasets has become commonplace. In such datasets, the data dimension can be extremely high, and the underlying data generation process can be unknown and highly…

Machine Learning · Statistics 2024-03-29 Yaxin Fang , Faming Liang

We consider the problems of classification and intrinsic dimension estimation on image data. A new subspace based classifier is proposed for supervised classification or intrinsic dimension estimation. The distribution of the data in each…

Computer Vision and Pattern Recognition · Computer Science 2020-02-11 Liang Liao , Stephen John Maybank

Estimating the dimensionality of the latent representation needed for prediction -- the task-relevant dimension -- is a difficult, largely unsolved problem with broad scientific applications. We cast it as an Information Bottleneck…

Machine Learning · Computer Science 2026-02-10 Paarth Gulati , Eslam Abdelaleem , Audrey Sederberg , Ilya Nemenman

Normality, in the colloquial sense, has historically been considered an aspirational trait, synonymous with ideality. The arithmetic average and, by extension, statistics including linear regression coefficients, have often been used to…

Methodology · Statistics 2023-12-27 Matthew J. Vowels

The curse of dimensionality in the realm of association rules is twofold. Firstly, we have the well known exponential increase in computational complexity with increasing item set size. Secondly, there is a \emph{related curse} concerned…

Artificial Intelligence · Computer Science 2018-05-16 Tom Hanika , Friedrich Martin Schneider , Gerd Stumme

High-dimensional dense embeddings have become central to modern Information Retrieval, but many dimensions are noisy or redundant. Recently proposed DIME (Dimension IMportance Estimation), provides query-dependent scores to identify…

Information Retrieval · Computer Science 2026-04-13 Giulio D'Erasmo , Cesare Campagnano , Antonio Mallia , Pierpaolo Brutti , Nicola Tonellotto , Fabrizio Silvestri

It has long been noticed that high dimension data exhibits strange patterns. This has been variously interpreted as either a "blessing" or a "curse", causing uncomfortable inconsistencies in the literature. We propose that these patterns…

Computer Vision and Pattern Recognition · Computer Science 2020-03-18 Wen-Yan Lin

Quantification of the number of variables needed to locally explain complex data is often the first step to better understanding it. Existing techniques from intrinsic dimension estimation leverage statistical models to glean this…

Machine Learning · Computer Science 2023-12-13 Eric Yeats , Cameron Darwin , Frank Liu , Hai Li

Modern datasets are characterized by a large number of features that may conceal complex dependency structures. To deal with this type of data, dimensionality reduction techniques are essential. Numerous dimensionality reduction methods…

Methodology · Statistics 2021-06-02 Francesco Denti , Diego Doimo , Alessandro Laio , Antonietta Mira

We address the problem of converting large-scale high-dimensional image data into binary codes so that approximate nearest-neighbor search over them can be efficiently performed. Different from most of the existing unsupervised approaches…

Computer Vision and Pattern Recognition · Computer Science 2015-12-02 Tsung-Yu Lin , Tsung-Wei Ke , Tyng-Luh Liu

A novel method for common and individual feature analysis from exceedingly large-scale data is proposed, in order to ensure the tractability of both the computation and storage and thus mitigate the curse of dimensionality, a major…

Signal Processing · Electrical Eng. & Systems 2017-11-03 Ilia Kisil , Giuseppe G. Calvi , Danilo P. Mandic

Missing values occur commonly in the multidimensional data warehouses. They may generate problems of usefulness of data since the analysis performed on a multidimensional data warehouse is through different dimensions with hierarchies where…

Databases · Computer Science 2021-10-05 Yuzhao Yang , Fatma Abdelhedi , Jérôme Darmont , Franck Ravat , Olivier Teste

Graphs are ubiquitous due to their flexibility in representing social and technological systems as networks of interacting elements. Graph representation learning methods, such as node embeddings, are powerful approaches to map nodes into a…

Machine Learning · Computer Science 2023-10-03 Simone Piaggesi , Megha Khosla , André Panisson , Avishek Anand