English
Related papers

Related papers: Information Distance Revisited

200 papers

We define the information threshold as the point of maximum curvature in the prior vs. posterior Bayesian curve, both of which are described as a function of the true positive and negative rates of the classification system in question. The…

Machine Learning · Statistics 2022-06-07 Jacques Balayla

Motivated by creating physical theories, formal languages $S$ with variables are considered and a kind of distance between elements of the languages is defined by the formula $d(x,y)= \ell(x \nabla y) - \ell(x) \wedge \ell(y)$, where $\ell$…

Information Theory · Computer Science 2023-11-17 Bernhard Burgstaller

In this paper we propose a Bayesian, information theoretic approach to dimensionality reduction. The approach is formulated as a variational principle on mutual information, and seamlessly addresses the notions of sufficiency, relevance,…

Data Analysis, Statistics and Probability · Physics 2007-05-23 David R. Wolf , Edward I. George

Cilibrasi and Vitanyi have demonstrated that it is possible to extract the meaning of words from the world-wide web. To achieve this, they rely on the number of webpages that are found through a Google search containing a given word and…

Computation and Language · Computer Science 2015-01-29 Bjørn Kjos-Hanssen , Alberto J. Evangelista

In this paper, we prove that two different observers don't equally measure the distance between two points A and B. For this, we introduce some postulates and obtain a new formula to show distance between A and B. In this formula, radius of…

General Physics · Physics 2007-05-23 M. Akbari , M. T. Darvishi

We establish concrete mathematical criteria to distinguish between different kinds of written storytelling, fictional and non-fictional. Specifically, we constructed a semantic network from both novels and news stories, with $N$ independent…

Computation and Language · Computer Science 2010-10-15 J. T. Stevanak , David M. Larue , Lincoln D. Carr

There are (at least) three approaches to quantifying information. The first, algorithmic information or Kolmogorov complexity, takes events as strings and, given a universal Turing machine, quantifies the information content of a string as…

Information Theory · Computer Science 2011-11-29 David Balduzzi

This short note introduces the harmonic indel distance (HID), a new distance between strings where the cost of an insertion or deletion is inversely proportional to the string length. We present a closed-form formula and show that the HID…

Discrete Mathematics · Computer Science 2024-12-13 Bob Pepin

How many bits of information are required to PAC learn a class of hypotheses of VC dimension $d$? The mathematical setting we follow is that of Bassily et al. (2018), where the value of interest is the mutual information…

Machine Learning · Computer Science 2018-04-20 Ido Nachum , Jonathan Shafer , Amir Yehudayoff

Ahlswede and Katona (1977) posed the following isodiametric problem in Hamming spaces: For every $n$ and $1\le M\le2^{n}$, determine the minimum average Hamming distance of binary codes with length $n$ and size $M$. Fu, Wei, and Yeung…

Combinatorics · Mathematics 2019-10-22 Lei Yu , Vincent Y. F. Tan

The Erd\H{o}s distance problem concerns the least number of distinct distances that can be determined by $N$ points in the plane. The integer lattice with $N$ points is known as \textit{near-optimal}, as it spans $\Theta(N/\sqrt{\log(N)})$…

In 1946 Erd\H os asked for the maximum number of unit distances, $u(n)$, among $n$ points in the plane. He showed that $u(n)> n^{1+c/\log\log n}$ and conjectured that this was the true magnitude. The best known upper bound is…

Combinatorics · Mathematics 2014-04-22 Ryan Schwartz , József Solymosi , Frank de Zeeuw

In the 2017 paper by Dougherty, Kim, Ozkaya, Sok, and Sol\'e about the linear programming bound for LCD codes the notion $\mathrm{LCD}[n,k]$ was defined for binary LCD $[n,k]$-codes. We find the formula for $\mathrm{LCD}[n,2]$.

Commutative Algebra · Mathematics 2019-09-04 Seth Gannon , Hamid Kulosman

In this paper we study the NP-Hard problem of maximizing the distance over an intersection of balls to a given point. We expand the results found in \cite{funcos1}, where the authors characterize the farthest in an intersection of balls…

Computational Geometry · Computer Science 2024-03-05 Beniamin Costandin , Marius Costandin

We investigate so-called localisable information of bipartite states and a parallel notion of information deficit. Localisable information is defined as the amount of information that can be concentrated by means of classical communication…

Quantum Physics · Physics 2015-06-26 Barbara Synak , Karol Horodecki , Michal Horodecki

The main contribution of this paper is to design an Information Retrieval (IR) technique based on Algorithmic Information Theory (using the Normalized Compression Distance- NCD), statistical techniques (outliers), and novel organization of…

Information Retrieval · Computer Science 2016-11-17 Rafael Martinez , Manuel Cebrian , Francisco de Borja Rodriguez , David Camacho

Let $A(n,d)$ be the maximum number of $0,1$ words of length $n$, any two having Hamming distance at least $d$. We prove $A(20,8)=256$, which implies that the quadruply shortened Golay code is optimal. Moreover, we show $A(18,6)\leq 673$,…

Combinatorics · Mathematics 2010-05-28 Dion C. Gijswijt , Hans D. Mittelmann , Alexander Schrijver

Let x, y be strings of equal length. The Hamming distance h(x,y) between x and y is the number of positions in which x and y differ. If x is a cyclic shift of y, we say x and y are conjugates. We consider f(x,y), the Hamming distance…

Combinatorics · Mathematics 2008-08-15 Jeffrey Shallit

This document reviews the definition of the kernel distance, providing a gentle introduction tailored to a reader with background in theoretical computer science, but limited exposure to technology more common to machine learning,…

Computational Geometry · Computer Science 2011-03-11 Jeff M. Phillips , Suresh Venkatasubramanian

We show that if X and Y are integers independently and uniformly distributed in the set {1, ..., N}, then the information lost in forming their product (which is given by the equivocation H(X,Y | XY)), is of order log log N. We also prove…

Probability · Mathematics 2007-05-23 Nicholas Pippenger