English
Related papers

Related papers: Measuring the Data

200 papers

High-dimensional data are ubiquitous in contemporary science and finding methods to compress them is one of the primary goals of machine learning. Given a dataset lying in a high-dimensional space (in principle hundreds to several thousands…

Machine Learning · Computer Science 2020-03-24 Vittorio Erba , Marco Gherardi , Pietro Rotondo

High dimensional data can have a surprising property: pairs of data points may be easily separated from each other, or even from arbitrary subsets, with high probability using just simple linear classifiers. However, this is more of a rule…

Machine Learning · Computer Science 2023-11-15 Oliver J. Sutton , Qinghua Zhou , Alexander N. Gorban , Ivan Y. Tyukin

The concept of dimension is essential to grasp the complexity of data. A naive approach to determine the dimension of a dataset is based on the number of attributes. More sophisticated methods derive a notion of intrinsic dimension (ID)…

Machine Learning · Computer Science 2023-04-18 Maximilian Stubbemann , Tom Hanika , Friedrich Martin Schneider

Manifold hypothesis states that data points in high-dimensional space actually lie in close vicinity of a manifold of much lower dimension. In many cases this hypothesis was empirically verified and used to enhance unsupervised and…

Quantum many-body systems are characterized by patterns of correlations that define highly-non trivial manifolds when interpreted as data structures. Physical properties of phases and phase transitions are typically retrieved via simple…

Statistical Mechanics · Physics 2021-09-01 Tiago Mendes-Santos , Adriano Angelone , Alex Rodriguez , Rosario Fazio , Marcello Dalmonte

Modern machine learning systems are increasingly trained on large amounts of data embedded in high-dimensional spaces. Often this is done without analyzing the structure of the dataset. In this work, we propose a framework to study the…

Machine Learning · Computer Science 2023-04-27 Carlos Hurtado , Sarath Shekkizhar , Javier Ruiz-Hidalgo , Antonio Ortega

Modern machine learning increasingly leverages the insight that high-dimensional data often lie near low-dimensional, non-linear manifolds, an idea known as the manifold hypothesis. By explicitly modeling the geometric structure of data…

Machine Learning · Computer Science 2026-03-02 Willem Diepeveen , Deanna Needell

Bigdata is a dataset of which size is beyond the ability of handling a valuable raw material that can be refined and distilled into valuable specific insights. Compact data is a method that optimizes the big dataset that gives best assets…

Databases · Computer Science 2020-12-29 Song-Kyoo , Kim

Analyzing large volumes of high-dimensional data is an issue of fundamental importance in data science, molecular simulations and beyond. Several approaches work on the assumption that the important content of a dataset belongs to a…

Machine Learning · Statistics 2018-03-20 Elena Facco , Maria d'Errico , Alex Rodriguez , Alessandro Laio

Quantification of the number of variables needed to locally explain complex data is often the first step to better understanding it. Existing techniques from intrinsic dimension estimation leverage statistical models to glean this…

Machine Learning · Computer Science 2023-12-13 Eric Yeats , Cameron Darwin , Frank Liu , Hai Li

In this work, we investigate Riemannian geometry based dimensionality reduction methods that respect the underlying manifold structure of the data. In particular, we focus on Principal Geodesic Analysis (PGA) as a nonlinear generalization…

Machine Learning · Computer Science 2026-02-06 Alaa El Ichi , Khalide Jbilou

Motivated by the manifold hypothesis, which states that data with a high extrinsic dimension may yet have a low intrinsic dimension, we develop refined statistical bounds for entropic optimal transport that are sensitive to the intrinsic…

Statistics Theory · Mathematics 2023-08-25 Austin J. Stromme

Dimensionality reduction is an effective method for learning high-dimensional data, which can provide better understanding of decision boundaries in human-readable low-dimensional subspace. Linear methods, such as principal component…

Machine Learning · Computer Science 2020-07-09 Koji Maruhashi , Heewon Park , Rui Yamaguchi , Satoru Miyano

Data augmentation is commonly used to encode invariances in learning methods. However, this process is often performed in an inefficient manner, as artificial examples are created by applying a number of transformations to all points in the…

Machine Learning · Computer Science 2019-03-04 Michael Kuchnik , Virginia Smith

Despite significant advances in the field of deep learning in applications to various fields, explaining the inner processes of deep learning models remains an important and open question. The purpose of this article is to describe and…

Machine Learning · Computer Science 2022-04-20 German Magai , Anton Ayzenberg

Big data often has emergent structure that exists at multiple levels of abstraction, which are useful for characterizing complex interactions and dynamics of the observations. Here, we consider multiple levels of abstraction via a…

Even the best scientific equipment can only partially observe reality. Recorded data is often lower-dimensional, e.g., two-dimensional pictures of the three-dimensional world. Combining data from multiple experiments then results in a…

Data Analysis, Statistics and Probability · Physics 2023-05-26 Michael Plainer , Felix Dietrich , Ioannis G. Kevrekidis

The intrinsic dimensionality refers to the ``true'' dimensionality of the data, as opposed to the dimensionality of the data representation. For example, when attributes are highly correlated, the intrinsic dimensionality can be much lower…

Machine Learning · Statistics 2020-11-30 Erik Thordsen , Erich Schubert

Topological methods can provide a way of proposing new metrics and methods of scrutinising data, that otherwise may be overlooked. In this work, a method of quantifying the shape of data, via a topic called topological data analysis will be…

Machine Learning · Statistics 2022-09-25 Tristan Gowdridge , Nikolaos Dervilis , Keith Worden

In the last decades the estimation of the intrinsic dimensionality of a dataset has gained considerable importance. Despite the great deal of research work devoted to this task, most of the proposed solutions prove to be unreliable when the…

Machine Learning · Computer Science 2012-06-19 Claudio Ceruti , Simone Bassis , Alessandro Rozza , Gabriele Lombardi , Elena Casiraghi , Paola Campadelli