English
Related papers

Related papers: The length of a typical Huffman codeword

200 papers

It is known that there are infinite words over finite alphabets with Abelian repetition threshold arbitrarily close to 1; however, the construction previously used involves huge alphabets. In this note we give a short cyclic morphism…

Combinatorics · Mathematics 2023-12-29 James D. Currie , Narad Rampersad

We consider a general class of super-additive scores measuring the similarity of two independent sequences of $n$ i.i.d. letters from a finite alphabet. Our object of interest is the mean score by letter $l_n$. By the subadditivity $l_n$ is…

Probability · Mathematics 2010-11-18 Juri Lember , Heinrich Matzinger , Felipe Torres

The penalty incurred by imposing a finite delay constraint in lossless source coding of a memoryless source is investigated. It is well known that for the so-called block-to-variable and variable-to-variable codes, the redundancy decays at…

Information Theory · Computer Science 2016-11-17 Ofer Shayevitz , Eado Meron , Meir Feder , Ram Zamir

It is argued that there are characteristic intervals associated with any particle that can be derived without reference to the speed of light $c$. Such intervals are inferred from zeros of wavefunctions which are solutions to the…

General Physics · Physics 2015-01-12 Mark D. Roberts

We evaluate the influence of different alphabet orderings on the Lyndon factorization of a string. Experiments with Pizza & Chili datasets show that for most alphabet reorderings, the number of Lyndon factors is usually small, and the…

Data Structures and Algorithms · Computer Science 2021-08-12 Marcelo K. Albertini , Felipe A. Louza

We show that the number of $t$-ary trees with path length equal to $p$ is $\exp(h(t^{-1})t\log t \frac{p}{\log p}(1+o(1)))$, where $\entropy(x){=}{-}x\log x {-}(1{-}x)\log (1{-}x)$ is the binary entropy function. Besides its intrinsic…

Discrete Mathematics · Computer Science 2007-07-16 Gadiel Seroussi

Language models (LMs) estimate a probability distribution over strings in a natural language; these distributions are crucial for computing perplexity and surprisal in linguistics research. While we are usually concerned with measuring…

Computation and Language · Computer Science 2024-10-15 Tiago Pimentel , Clara Meister

The halting probabilities of universal prefix-free machines are universal for the class of reals with computably enumerable left cut (also known as left-c.e. reals), and coincide with the Martin-Loef random elements of this class. We study…

Computational Complexity · Computer Science 2017-05-22 George Barmpalias , Andrew Lewis-Pye

We simply construct a quantum universal variable-length source code in which, independent of information source, both of the average error and the probability that the coding rate is greater than the entropy rate $H(rho_p)$, tend to 0. If…

Quantum Physics · Physics 2007-05-23 Masahito Hayashi , Keiji Matsumoto

Given two random finite sequences from $[k]^n$ such that a prefix of the first sequence is a suffix of the second, we examine the length of their longest common subsequence. If $\ell$ is the length of the overlap, we prove that the expected…

Probability · Mathematics 2018-03-12 Boris Bukh , Raymond Hogenson

A $c$-short program for a string $x$ is a description of $x$ of length at most $C(x) + c$, where $C(x)$ is the Kolmogorov complexity of $x$. We show that there exists a randomized algorithm that constructs a list of $n$ elements that…

Computational Complexity · Computer Science 2015-01-21 Bruno Bauwens , Marius Zimand

When an i.i.d.\ sequence of letters is cut into words according to i.i.d.\ renewal times, an i.i.d.\ sequence of words is obtained. In the \emph{annealed} LDP (large deviation principle) for the empirical process of words, the rate function…

Probability · Mathematics 2013-11-22 Frank den Hollander , Julien Poisat

How hard is it guess a password? Massey showed that that the Shannon entropy of the distribution from which the password is selected is a lower bound on the expected number of guesses, but one which is not tight in general. In a series of…

Information Theory · Computer Science 2013-02-12 Mark M. Christiansen , Ken R. Duffy

Huffman Compression, also known as Huffman Coding, is one of many compression techniques in use today. The two important features of Huffman coding are instantaneousness that is the codes can be interpreted as soon as they are received and…

Information Theory · Computer Science 2013-02-22 A. S. Tolba , M. Z. Rashad , M. A. El-Dosuky

Let $\varepsilon>0$ and, for an odd prime $p$, set $$ S_\ell(p):=\sum_{n\le \ell}\left(\frac{n}{p}\right). $$ Define the first-passage time $$ f_\varepsilon(p):=\min\{\ell\ge 1:\ S_\ell(p)<\varepsilon\ell\}. $$ We prove that there exists a…

Number Theory · Mathematics 2026-01-21 Quanyu Tang , Hao Zhang

Given a collection of strings, each with an associated probability of occurrence, the guesswork of each of them is their position in a list ordered from most likely to least likely, breaking ties arbitrarily. Guesswork is central to several…

Information Theory · Computer Science 2019-08-12 Ahmad Beirami , Robert Calderbank , Mark Christiansen , Ken Duffy , Muriel Médard

Given a probability distribution over a set of n words to be transmitted, the Huffman Coding problem is to find a minimal-cost prefix free code for transmitting those words. The basic Huffman coding problem can be solved in O(n log n) time…

Data Structures and Algorithms · Computer Science 2008-09-29 Mordecai Golin , Xiaoming Xu , Jiajin Yu

We study the membership problem to context-free languages L (CFLs) on probabilistic words, that specify for each position a probability distribution on the letters (assuming independence across positions). Our task is to compute, given a…

Formal Languages and Automata Theory · Computer Science 2025-10-10 Antoine Amarilli , Mikaël Monet , Paul Raphaël , Sylvain Salvati

We define the ``shift-match number'' for a binary string and we compute the probability of occurrence of a given string as a subsequence in longer strings in terms of its shift-match number. We thus prove that the string matching…

Genomics · Quantitative Biology 2007-05-23 A. H. Bilge , A. Erzan , D. Balcan

The combined universal probability $\mathbf{m}(D)$ of strings $x$ in sets $D$ is close to max $\mathbf{m}(x)$ over $x$ in $D$: their logs differ by at most $D$'s information $\mathbf{I}(D:\mathcal{H})$ about the halting sequence…

Computational Complexity · Computer Science 2023-09-12 Samuel Epstein