Local Decodability of the Burrows-Wheeler Transform
Abstract
The Burrows-Wheeler Transform (BWT) is among the most influential discoveries in text compression and DNA storage. It is a reversible preprocessing step that rearranges an -letter string into runs of identical characters (by exploiting context regularities), resulting in highly compressible strings, and is the basis of the \texttt{bzip} compression program. Alas, the decoding process of BWT is inherently sequential and requires time even to retrieve a \emph{single} character. We study the succinct data structure problem of locally decoding short substrings of a given text under its \emph{compressed} BWT, i.e., with small additive redundancy over the \emph{Move-To-Front} (\texttt{bzip}) compression. The celebrated BWT-based FM-index (FOCS '00), as well as other related literature, yield a trade-off of bits, when a single character is to be decoded in time. We give a near-quadratic improvement . As a by-product, we obtain an \emph{exponential} (in ) improvement on the redundancy of the FM-index for counting pattern-matches on compressed text. In the interesting regime where the text compresses to bits, these results provide an \emph{overall} space reduction. For the local decoding problem of BWT, we also prove an cell-probe lower bound for "symmetric" data structures. We achieve our main result by designing a compressed partial-sums (Rank) data structure over BWT. The key component is a \emph{locally-decodable} Move-to-Front (MTF) code: with only extra bits per block of length , the decoding time of a single character can be decreased from to . This result is of independent interest in algorithmic information theory.
Keywords
Cite
@article{arxiv.1808.03978,
title = {Local Decodability of the Burrows-Wheeler Transform},
author = {Sandip Sinha and Omri Weinstein},
journal= {arXiv preprint arXiv:1808.03978},
year = {2018}
}
Comments
The following two technical typos were fixed: (1) On page 2, following Theorem 1, the decoding time of a contiguous substring of size $\ell$ was corrected from $O(t + \ell)$ to $O(t + \ell \cdot \lg t)$. (2) In the statement of Theorem 2, the query time to count occurrences of patterns of length $\ell$ was corrected to $O(t \ell)$, independent of the number of occurrences