Prefix-Free Parsing for Building Big BWTs
Abstract
High-throughput sequencing technologies have led to explosive growth of genomic databases; one of which will soon reach hundreds of terabytes. For many applications we want to build and store indexes of these databases but constructing such indexes is a challenge. Fortunately, many of these genomic databases are highly-repetitive---a characteristic that can be exploited to ease the computation of the Burrows-Wheeler Transform (BWT), which underlies many popular indexes. In this paper, we introduce a preprocessing algorithm, referred to as {\em prefix-free parsing}, that takes a text as input, and in one-pass generates a dictionary and a parse of with the property that the BWT of can be constructed from and using workspace proportional to their total size and -time. Our experiments show that and are significantly smaller than in practice, and thus, can fit in a reasonable internal memory even when is very large. In particular, we show that with prefix-free parsing we can build an 131-megabyte run-length compressed FM-index (restricted to support only counting and not locating) for 1000 copies of human chromosome 19 in 2 hours using 21 gigabytes of memory, suggesting that we can build a 6.73 gigabyte index for 1000 complete human-genome haplotypes in approximately 102 hours using about 1 terabyte of memory.
Keywords
Cite
@article{arxiv.1803.11245,
title = {Prefix-Free Parsing for Building Big BWTs},
author = {Christina Boucher and Travis Gagie and Alan Kuhnle and Ben Langmead and Giovanni Manzini and Taher Mun},
journal= {arXiv preprint arXiv:1803.11245},
year = {2018}
}
Comments
Preliminary version appeared at WABI '18; full version submitted to a journal