English
Related papers

Related papers: Ultrametric embedding: application to data fingerp…

200 papers

We present new findings in regard to data analysis in very high dimensional spaces. We use dimensionalities up to around one million. A particular benefit of Correspondence Analysis is its suitability for carrying out an orthonormal…

Machine Learning · Statistics 2015-12-15 Fionn Murtagh

Although distance measures are used in many machine learning algorithms, the literature on the context-independent selection and evaluation of distance measures is limited in the sense that prior knowledge is used. In cluster analysis,…

Machine Learning · Computer Science 2021-08-24 Michael C. Thrun

Motivated by the fact that 2-dimensional data have become popularly used in many applications without being much considered its integrity checking. We introduce the problem of detecting a corrupted area in a 2-dimensional space, and…

Data Structures and Algorithms · Computer Science 2014-04-08 YounSun Cho

Time series constitute a challenging data type for machine learning algorithms, due to their highly variable lengths and sparse labeling in practice. In this paper, we tackle this challenge by proposing an unsupervised method to learn…

Machine Learning · Computer Science 2020-01-06 Jean-Yves Franceschi , Aymeric Dieuleveut , Martin Jaggi

Temporal sequences of satellite images constitute a highly valuable and abundant resource for analyzing regions of interest. However, the automatic acquisition of knowledge on a large scale is a challenging task due to different factors…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Carlos Echegoyen , Aritz Pérez , Guzmán Santafé , Unai Pérez-Goya , María Dolores Ugarte

Recently, many works have focused on the characterization of non-linear dimensionality reduction methods obtained by quantizing linear embeddings, e.g., to reach fast processing time, efficient data compression procedures, novel…

Information Theory · Computer Science 2016-12-30 Laurent Jacques , Valerio Cambareri

There has been an intense recent activity in embedding of very high dimensional and nonlinear data structures, much of it in the data science and machine learning literature. We survey this activity in four parts. In the first part we cover…

Machine Learning · Statistics 2022-09-01 Dag Tjøstheim , Martin Jullum , Anders Løland

We describe many vantage points on the Baire metric and its use in clustering data, or its use in preprocessing and structuring data in order to support search and retrieval operations. In some cases, we proceed directly to clusters and do…

Machine Learning · Statistics 2016-09-20 Fionn Murtagh , Pedro Contreras

We give a practical random mapping that takes any set of documents represented as vectors in Euclidean space and then maps them to a sparse subset of the Hamming cube while retaining ordering of inter-vector inner products. Once represented…

Information Retrieval · Computer Science 2015-07-22 Roger Donaldson , Arijit Gupta , Yaniv Plan , Thomas Reimer

Teramoto et al. defined a new measure called the gap ratio that measures the uniformity of a finite point set sampled from $\cal S$, a bounded subset of $\mathbb{R}^2$. We generalize this definition of measure over all metric spaces by…

Computational Geometry · Computer Science 2015-12-07 Arijit Bishnu , Sameer Desai , Arijit Ghosh , Mayank Goswami , Subhabrata Paul

Text indexing is a fundamental and well-studied problem. Classic solutions either replace the original text with a compressed representation, e.g., the FM-index and its variants, or keep it uncompressed but attach some redundancy - an index…

Data Structures and Algorithms · Computer Science 2026-02-05 Lorraine A. K. Ayad , Gabriele Fici , Ragnar Groot Koerkamp , Grigorios Loukides , Rob Patro , Giulio Ermanno Pibiri , Solon P. Pissis

Quantum embedding methods have become a powerful tool to overcome deficiencies of traditional quantum modelling in materials science. However, while these are systematically improvable in principle, in practice it is rarely possible to…

Strongly Correlated Electrons · Physics 2022-11-02 Max Nusspickel , George H. Booth

Entropy is the measure of uncertainty in any data and is adopted for maximisation of mutual information in many remote sensing operations. The availability of wide entropy variations motivated us for an investigation over the suitability…

Computer Vision and Pattern Recognition · Computer Science 2014-05-26 S. K. Katiyar , P. V. Arun

Important properties of a quantum system are not directly measurable, but they can be disclosed by how fast the system changes under controlled perturbations. In particular, asymmetry and entanglement can be verified by reconstructing the…

As the proliferation of high-throughput approaches in materials science is increasing the wealth of data in the field, the gap between accumulated-information and derived-knowledge widens. We address the issue of scientific discovery in…

We present a general approach to the study of the local distribution of measures on Euclidean spaces, based on local entropy averages. As concrete applications, we unify, generalize, and simplify a number of recent results on local…

Classical Analysis and ODEs · Mathematics 2015-02-03 Tuomas Sahlsten , Pablo Shmerkin , Ville Suomala

Similarity query is the family of queries based on some similarity metrics. Unlike the traditional database queries which are mostly based on value equality, similarity queries aim to find targets "similar enough to" the given data objects,…

Databases · Computer Science 2022-04-19 Yifan Wang

Imaging spectrometers measure electromagnetic energy scattered in their instantaneous field view in hundreds or thousands of spectral channels with higher spectral resolution than multispectral cameras. Imaging spectrometers are therefore…

Data Analysis, Statistics and Probability · Physics 2012-04-25 José M. Bioucas-Dias , Antonio Plaza , Nicolas Dobigeon , Mario Parente , Qian Du , Paul Gader , Jocelyn Chanussot

DBSCAN* and HDBSCAN* are well established density based clustering algorithms. However, obtaining the clusters of very large datasets is infeasible, limiting their use in real world applications. By exploiting the geometry of Euclidean…

Machine Learning · Computer Science 2022-03-16 A. L. Garcia-Pulido , K. P. Samardzhiev

We prove optimal bounds for the convergence rate of ordinal embedding (also known as non-metric multidimensional scaling) in the 1-dimensional case. The examples witnessing optimality of our bounds arise from a result in additive number…

Statistics Theory · Mathematics 2019-05-01 Jordan S. Ellenberg , Lalit Jain