English

Compression of data streams down to their information content

Information Theory 2019-01-23 v4 math.IT

Abstract

According to Kolmogorov complexity, every finite binary string is compressible to a shortest code -- its information content -- from which it is effectively recoverable. We investigate the extent to which this holds for infinite binary sequences (streams). We devise a new coding method which uniformly codes every stream XX into an algorithmically random stream YY, in such a way that the first nn bits of XX are recoverable from the first I(Xn)I(X\upharpoonright_n) bits of YY, where II is any partial computable information content measure which is defined on all prefixes of XX, and where XnX\upharpoonright_n is the initial segment of XX of length nn. As a consequence, if gg is any computable upper bound on the initial segment prefix-free complexity of XX, then XX is computable from an algorithmically random YY with oracle-use at most gg. Alternatively (making no use of such a computable bound gg) one can achieve an oracle-use bounded above by K(Xn)+lognK(X\upharpoonright_n)+\log n. This provides a strong analogue of Shannon's source coding theorem for algorithmic information theory.

Keywords

Cite

@article{arxiv.1710.02092,
  title  = {Compression of data streams down to their information content},
  author = {George Barmpalias and Andrew Lewis-Pye},
  journal= {arXiv preprint arXiv:1710.02092},
  year   = {2019}
}
R2 v1 2026-06-22T22:04:51.958Z