English
Related papers

Related papers: Properties of Algorithmic Information Distance

200 papers

The normalized edit distance is one of the distances derived from the edit distance. It is useful in some applications because it takes into account the lengths of the two strings compared. The normalized edit distance is not defined in…

Neural and Evolutionary Computing · Computer Science 2013-12-09 Muhammad Marwan Muhammad Fuad

A nonlinear generalisation of Schrodinger's equation is obtained using information-theoretic arguments. The nonlinearities are controlled by an intrinsic length scale and involve derivatives to all orders thus making the equation mildly…

High Energy Physics - Theory · Physics 2009-11-10 Rajesh R. Parwani

A configuration p in r-dimensional Euclidean space is a finite collection of labeled points p^1,p^2,...,p^n in R^r that affinely span R^r. Each configuration p defines a Euclidean distance matrix D_p = (d_ij) = (||p^i-p^j||^2), where ||.||…

Metric Geometry · Mathematics 2012-01-17 A. Y. Alfakih

Embedding data into vector spaces is a very popular strategy of pattern recognition methods. When distances between embeddings are quantized, performance metrics become ambiguous. In this paper, we present an analysis of the ambiguity…

Computer Vision and Pattern Recognition · Computer Science 2019-02-21 Anguelos Nicolaou , Sounak Dey , Vincent Christlein , Andreas Maier , Dimosthenis Karatzas

We study the complexity of the classic capacitated k-median and k-means problems parameterized by the number of centers, k. These problems are notoriously difficult since the best known approximation bound for high dimensional Euclidean…

Data Structures and Algorithms · Computer Science 2022-08-31 Vincent Cohen-Addad , Jason Li

Many network information theory problems face the similar difficulty of single-letterization. We argue that this is due to the lack of a geometric structure on the space of probability distribution. In this paper, we develop such a…

Information Theory · Computer Science 2014-06-12 Shao-Lun Huang , Lizhong Zheng

Localization of a set of nodes is an important and a thoroughly researched problem in robotics and sensor networks. This paper is concerned with the theory of localization from inner-angle measurements. We focus on the challenging case…

Robotics · Computer Science 2020-05-12 Frederike Dümbgen , Majed El Helou , Adam Scholefield

The Euclidean distance geometry problem arises in a wide variety of applications, from determining molecular conformations in computational chemistry to localization in sensor networks. When the distance information is incomplete, the…

Information Theory · Computer Science 2018-10-30 Abiy Tasissa , Rongjie Lai

A fundamental question in data analysis, machine learning and signal processing is how to compare between data points. The choice of the distance metric is specifically challenging for high-dimensional data sets, where the problem of…

Machine Learning · Statistics 2017-08-15 Almog Lahav , Ronen Talmon , Yuval Kluger

It is known that the normalized algorithmic information distance $N$ is not computable and not semicomputable. We show that for all $\epsilon < 1/2$, there exist no semicomputable functions that differ from $N$ by at most~$\epsilon$.…

Information Theory · Computer Science 2020-02-18 Bruno Bauwens , Ilya Blinnikov

An appropriate distance metric is crucial for categorical data clustering, as the distance between categorical data cannot be directly calculated. However, the distances between attribute values usually vary in different clusters induced by…

Machine Learning · Computer Science 2026-03-09 Taixi Chen , Yiu-ming Cheung , Yiqun Zhang

The manifold hypothesis suggests that the generalization performance of machine learning methods improves significantly when the intrinsic dimension of the input distribution's support is low. In the context of KRR, we investigate two…

Machine Learning · Computer Science 2026-01-23 Rustem Takhanov

Many scientific fields study data with an underlying structure that is a non-Euclidean space. Some examples include social networks in computational social sciences, sensor networks in communications, functional networks in brain imaging,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-02 Michael M. Bronstein , Joan Bruna , Yann LeCun , Arthur Szlam , Pierre Vandergheynst

We revisit Pollard's classical result on consistency for $k$-means clustering in Euclidean space, with a focus on extensions in two directions: first, to problems where the data may come from interesting geometric settings (e.g., Riemannian…

Statistics Theory · Mathematics 2025-07-01 Adam Quinn Jaffe

Normalized Compression Distance (NCD) is a popular tool that uses compression algorithms to cluster and classify data in a wide range of applications. Existing discussions of NCD's theoretical merit rely on certain theoretical properties of…

Cryptography and Security · Computer Science 2015-09-03 Rebecca Schuller Borbely

In a recent paper the author proved a theorem to the effect that the matrix of normalized Euclidean distances on the set of specially distributed random points in the $n$-dimensional Euclidean space $\mathbb R^{n}$ with independent…

Mathematical Physics · Physics 2015-09-07 A. P. Zubarev

We prove that every online learnable class of functions of Littlestone dimension $d$ admits a learning algorithm with finite information complexity. Towards this end, we use the notion of a globally stable algorithm. Generally, the…

Machine Learning · Computer Science 2022-06-28 Aditya Pradeep , Ido Nachum , Michael Gastpar

The Euclidean k-means problem is arguably the most widely-studied clustering problem in machine learning. While the k-means objective is NP-hard in the worst-case, practitioners have enjoyed remarkable success in applying heuristics like…

Machine Learning · Computer Science 2017-12-05 Abhratanu Dutta , Aravindan Vijayaraghavan , Alex Wang

We revisit extending the Kolmogorov-Smirnov distance between probability distributions to the multidimensional setting and make new arguments about the proper way to approach this generalization. Our proposed formulation maximizes the…

Computation · Statistics 2025-04-16 Peter Matthew Jacobs , Foad Namjoo , Jeff M. Phillips

K-means defines one of the most employed centroid-based clustering algorithms with performances tied to the data's embedding. Intricate data embeddings have been designed to push $K$-means performances at the cost of reduced theoretical…

Machine Learning · Computer Science 2022-02-17 Romain Cosentino , Randall Balestriero , Yanis Bahroun , Anirvan Sengupta , Richard Baraniuk , Behnaam Aazhang
‹ Prev 1 3 4 5 6 7 10 Next ›