English

Average Size of a Suffix Tree for Markov Sources

Data Structures and Algorithms 2016-05-10 v1

Abstract

We study a suffix tree built from a sequence generated by a Markovian source. Such sources are more realistic probabilistic models for text generation, data compression, molecular applications, and so forth. We prove that the average size of such a suffix tree is asymptotically equivalent to the average size of a trie built over nn independent sequences from the same Markovian source. This equivalence is only known for memoryless sources. We then derive a formula for the size of a trie under Markovian model to complete the analysis for suffix trees. We accomplish our goal by applying some novel techniques of analytic combinatorics on words also known as analytic pattern matching.

Cite

@article{arxiv.1605.02123,
  title  = {Average Size of a Suffix Tree for Markov Sources},
  author = {Philippe Jacquet and Wojciech Szpankowski},
  journal= {arXiv preprint arXiv:1605.02123},
  year   = {2016}
}

Comments

AofA 2016 Conference

R2 v1 2026-06-22T13:55:19.071Z