English
Related papers

Related papers: Classifying token frequencies using angular Minkow…

200 papers

The mutual information is bounded from above by a decreasing affine function of the square of the distance between the input distribution and the set of all capacity-achieving input distributions $\Pi_{\mathcal{A}}$, on small enough…

Information Theory · Computer Science 2025-04-24 Barış Nakiboğlu , Hao-Chung Cheng

Minkowski Functionals are a powerful tool for analyzing large scale structure, in particular if the distribution of matter is highly non-Gaussian, as it is in models in which cosmic strings contribute to structure formation. Here we apply…

Cosmology and Nongalactic Astrophysics · Physics 2015-03-19 Evan McDonough , Robert H. Brandenberger

The theory of p-adic fractal strings and their complex dimensions was developed by the first two authors in [17, 18, 19], particularly in the self-similar case, in parallel with its archimedean (or real) counterpart developed by the first…

Metric Geometry · Mathematics 2014-03-27 Michel L. Lapidus , Lu Hung , Machiel van Frankenhuijsen

The correlation distance quantifies the statistical independence of two classical or quantum systems, via the distance from their joint state to the product of the marginal states. Tight lower bounds are given for the mutual information…

Quantum Physics · Physics 2013-09-23 Michael J. W. Hall

Memory-based collaborative filtering methods like user or item k-nearest neighbors (kNN) are a simple yet effective solution to the recommendation problem. The backbone of these methods is the estimation of the empirical similarity between…

Information Retrieval · Computer Science 2019-05-20 Farhan Khawar , Nevin L. Zhang

Topological Data Analysis methods can be useful for classification and clustering tasks in many different fields as they can provide two dimensional persistence diagrams that summarize important information about the shape of potentially…

Quantum Physics · Physics 2024-09-02 Bernardo Ameneyro , Rebekah Herrman , George Siopsis , Vasileios Maroulas

The k-nearest neighbors (k-NN) is a basic machine learning (ML) algorithm, and several quantum versions of it, employing different distance metrics, have been presented in the last few years. Although the Euclidean distance is one of the…

Emerging Technologies · Computer Science 2024-04-25 Enrico Zardini , Enrico Blanzieri , Davide Pastorello

In this paper, we apply an efficient top-$k$ shortest distance routing algorithm to the link prediction problem and test its efficacy. We compare the results with other base line and state-of-the-art methods as well as with the shortest…

Social and Information Networks · Computer Science 2017-05-09 Andrei Lebedev , JooYoung Lee , Victor Rivera , Manuel Mazzara

This paper studies a discrepancy-sensitive approach to dynamic fractional cascading. We provide an efficient data structure for dominated maxima searching in a dynamic set of points in the plane, which in turn leads to an efficient dynamic…

Data Structures and Algorithms · Computer Science 2009-04-30 Mikhail J. Atallah , Marina Blanton , Michael T. Goodrich , Stanislas Polu

We survey the emerging area of compression-based, parameter-free, similarity distance measures useful in data-mining, pattern recognition, learning and automatic semantics extraction. Given a family of distances on a set of objects, a…

Computer Vision and Pattern Recognition · Computer Science 2007-05-23 Rudi Cilibrasi , Paul Vitanyi

In the realm of machine learning, the KNN classification algorithm is widely recognized for its simplicity and efficiency. However, its sensitivity to the K value poses challenges, especially with small sample sizes or outliers, impacting…

Machine Learning · Computer Science 2024-05-29 Junzhuo Chen , Zhixin Lu , Shitong Kang

In the study of Euclidean lattices, the product of the successive minima is bounded from above and below by explicit quantities. This result is known as Minkowski's second theorem, and can be refined to include Hermite's constant in the…

Number Theory · Mathematics 2025-07-22 Mathieu Dutour

Probabilistic k-nearest neighbour (PKNN) classification has been introduced to improve the performance of original k-nearest neighbour (KNN) classification algorithm by explicitly modelling uncertainty in the classification of each feature…

Machine Learning · Computer Science 2013-05-07 Ji Won Yoon , Nial Friel

We survey a new area of parameter-free similarity distance measures useful in data-mining, pattern recognition, learning and automatic semantics extraction. Given a family of distances on a set of objects, a distance is universal up to a…

Information Retrieval · Computer Science 2007-05-23 Paul Vitanyi

Compositionality in language refers to how much the meaning of some phrase can be decomposed into the meaning of its constituents and the way these constituents are combined. Based on the premise that substitution by synonyms is…

Computation and Language · Computer Science 2017-03-13 Christina Lioma , Niels Dalum Hansen

Distance metrics and their nonlinear variant play a crucial role in machine learning based real-world problem solving. We demonstrated how Euclidean and cosine distance measures differ not only theoretically but also in real-world medical…

Machine Learning · Computer Science 2021-02-25 Der-Chen Chang , Ophir Frieder , Chi-Feng Hung , Hao-Ren Yao

Locality Sensitive Filters are known for offering a quasi-linear space data structure with rigorous guarantees for the Approximate Near Neighbor search (ANN) problem. Building on Locality Sensitive Filters, we derive a simple data structure…

Data Structures and Algorithms · Computer Science 2025-05-05 Martin Aumüller , Fabrizio Boninsegna , Francesco Silvestri

We propose a novel method to determine the dissimilarity between subjects for functional data clustering. Spline smoothing or interpolation is common to deal with data of such type. Instead of estimating the best-representing curve for each…

Methodology · Statistics 2021-03-23 ShengLi Tzeng , Christian Hennig , Yu-Fen Li , Chien-Ju Lin

We extend the notion of the distance to a measure from Euclidean space to probability measures on general metric spaces as a way to do topological data analysis in a way that is robust to noise and outliers. We then give an efficient way to…

Computational Geometry · Computer Science 2014-10-09 Mickael Buchet , Frederic Chazal , Steve Y. Oudot , Donald R. Sheehy

Minkowski tensors, also known as tensor valuations, provide robust $n$-point information for a wide range of random spatial structures. Local estimators for point clouds, e.g., representing voxelized data, however, are unavoidably biased…

Statistics Theory · Mathematics 2026-04-06 Daniel Hug , Michael A. Klatt , Dominik Pabst