The landscape of compressibility measures for two-dimensional data
Abstract
In this paper we extend to two-dimensional data two recently introduced one-dimensional compressibility measures: the measure defined in terms of the smallest string attractor, and the measure defined in terms of the number of distinct substrings of the input string. Concretely, we introduce the two-dimensional measures and , as natural generalizations of and , and we initiate the study of their properties. Among other things, we prove that is monotone and can be computed in linear time, and we show that, although it is still true that , the gap between the two measures can be and therefore asymptotically larger than the gap between and . To complete the scenario of two-dimensional compressibility measures, we introduce the measure which generalizes to two dimensions the notion of optimal parsing. We prove that, somewhat surprisingly, the relationship between and is significantly different than in the one-dimensional case. As an application of our results we provide the first analysis of the space usage of the two-dimensional block tree introduced in [Brisaboa et al., Two-dimensional block trees, The computer Journal, 2024]. Our analysis shows that the space usage can be bounded in terms of both and . Finally, using insights from our analysis, we design the first linear time and space algorithm for constructing the two-dimensional block tree for arbitrary matrices.
Cite
@article{arxiv.2307.02629,
title = {The landscape of compressibility measures for two-dimensional data},
author = {Lorenzo Carfagna and Giovanni Manzini},
journal= {arXiv preprint arXiv:2307.02629},
year = {2024}
}