Related papers: Large Alphabets and Incompressibility
We study the embedding $\text{id}: \ell_p^b(\ell_q^d) \to \ell_r^b(\ell_u^d)$ and prove matching bounds for the entropy numbers $e_k(\text{id})$ provided that $0<p<r\leq \infty$ and $0<q\leq u\leq \infty$. Based on this finding, we…
We investigate the longest common substring problem for encoded sequences and its asymptotic behaviour. The main result is a strong law of large numbers for a re-scaled version of this quantity, which presents an explicit relation with the…
Since the introduction of the Kolmogorov complexity of binary sequences in the 1960s, there have been significant advancements in the topic of complexity measures for randomness assessment, which are of fundamental importance in theoretical…
In this paper we develop the following general approach. We study asymptotic behavior of the entropy numbers not for an individual smoothness class, how it is usually done, but for the collection of classes, which are defined by integral…
We study practical approximations to Kolmogorov prefix complexity (K) using IMP2, a high-level programming language. Our focus is on investigating the interpreter optimality for this language as the reference machine for the Coding Theorem…
We consider the task of estimating a conditional density using i.i.d. samples from a joint distribution, which is a fundamental problem with applications in both classification and uncertainty quantification for regression. For joint…
Consider the problem of estimating the Shannon entropy of a distribution over $k$ elements from $n$ independent samples. We show that the minimax mean-square error is within universal multiplicative constant factors of $$\Big(\frac{k }{n…
Several problems in machine learning, statistics, and other fields rely on computing eigenvectors. For large scale problems, the computation of these eigenvectors is typically performed via iterative schemes such as subspace iteration or…
In many high-impact applications, it is important to ensure the quality of output of a machine learning algorithm as well as its reliability in comparison with the complexity of the algorithm used. In this paper, we have initiated a…
Kolmogorov-Chaitin complexity has long been believed to be impossible to approximate when it comes to short sequences (e.g. of length 5-50). However, with the newly developed \emph{coding theorem method} the complexity of strings of length…
We study concentration inequalities for the Kullback--Leibler (KL) divergence between the empirical distribution and the true distribution. Applying a recursion technique, we improve over the method of types bound uniformly in all regimes…
We analyze the algorithm in [Holub, 2009], which decides whether a given word is a fixed point of a nontrivial morphism. We show that it can be implemented to have complexity in O(mn), where n is the length of the word and m the size of the…
We present two new methods for estimating the order (memory depth) of a finite alphabet Markov chain from observation of a sample path. One method is based on entropy estimation via recurrence times of patterns, and the other relies on a…
The problem of positive Kolmogorov-Sinai entropy of the Chirikov-Standard map with respect to the invariant Lebesgue measure on the two-dimensional is open. In 1999, we believed to have a proof that the entropy can be bounded below. This…
Interacting random field of probabilities links Kolmogorov law 0-1 and Bayesian probabilities observing Markov diffusion process under Yes-No actions of random impulse. These objective probabilities measure virtual probing impulses…
In this paper we give a detailed analysis of deterministic and randomized algorithms that enumerate any number of irreducible polynomials of degree $n$ over a finite field and their roots in the extension field in quasilinear where $N=n^2$…
We investigate the generalization properties of dense text embeddings when the embedding backbone is a large language model (LLM) versus when it is a non-LLM encoder, and we study the extent to which spherical linear interpolation (SLERP)…
Many algorithms are specified with respect to a fixed but unspecified parameter. Examples of this are especially common in cryptography, where protocols often feature a security parameter such as the bit length of a secret key. Our aim is…
Building upon the work of Chebyshev, Shannon and Kontoyiannis, it may be demonstrated that Chebyshev's asymptotic result: \begin{equation} \ln N \sim \sum_{p \leq N} \frac{1}{p} \cdot \ln p \end{equation} has a natural information-theoretic…
We discuss inequalities holding between the vocabulary size, i.e., the number of distinct nonterminal symbols in a grammar-based compression for a string, and the excess length of the respective universal code, i.e., the code-based analog…