R-enum Revisited: Speedup and Extension for Context-Sensitive Repeats and Net Frequencies
Abstract
Nishimoto and Tabei [CPM, 2021] proposed r-enum, an algorithm to enumerate various characteristic substrings, including maximal repeats, in a string of length in words of compressed working space, where is the number of runs in the Burrows-Wheeler transform (BWT) of . Given the run-length encoded BWT (RLBWT) of , r-enum runs in time in addition to the time linear to the number of output strings, where is the word size. In this paper, we first improve the term to . We next extend r-enum to compute other context-sensitive repeats such as near-supermaximal repeats (NSMRs) and supermaximal repeats, as well as the context diversity for every maximal repeat in the same complexities. Furthermore, we study net occurrences: An occurrence of a repeat is called a net occurrence if it is not covered by another repeat, and the net frequency of a repeat is the number of its net occurrences. With this terminology, an NSMR is a repeat with a positive net frequency. Given the RLBWT of , we show how to compute the set of all NSMRs in together with their net frequency/occurrences in time and space. We also show that an -space data structure can be built from the RLBWT to compute the net frequency/occurrences of any pattern in optimal time. The data structure is built in space and in time with high probability or deterministic time, where is the alphabet size of . To achieve this, we prove that the total number of net occurrences is less than . With the duality between net occurrences and \emph{minimal unique substrings (MUSs)}, we get a new upper bound of the number of MUSs in , which may be of independent interest.
Cite
@article{arxiv.2511.11057,
title = {R-enum Revisited: Speedup and Extension for Context-Sensitive Repeats and Net Frequencies},
author = {Kotaro Kimura and Tomohiro I},
journal= {arXiv preprint arXiv:2511.11057},
year = {2026}
}
Comments
accepted to CPM2026