English

Tight Bounds for the Number of Absent Subsequences

Formal Languages and Automata Theory 2025-09-01 v2 Combinatorics

Abstract

A {\em subsequence} of a word ww is a word uu that can be obtained by deleting some letters from ww while maintaining the relative order of the remaining letters, e.g., lala\mathtt{lala} is a subsequence of alfalfa\mathtt{alfalfa}. A word, over some alphabet Σ\Sigma, which has all possible words of length ι\iota over Σ\Sigma as subsequences is called ι\iota-universal, and the largest ι\iota for which this holds is called the universality index of ww, and denoted ι(w)\iota(w). Moreover, words that are not subsequences of ww are called absent subsequences (AS) of ww, and their investigation was started in (Kosche et al., 2022). In this paper, we present tight bounds on the number of AS of a given length kk among all words with the same universality index ι\iota. For both the lower and upper bound, we construct words that have, respectively, a minimal and maximal number of absent subsequences of the respective length kk, and, in the case of the lower bound, we provide the exact number of missing subsequences as a closed form. Finally, we present efficient enumeration algorithms for the set of subsequences of given length of a word: we give a novel, optimal enumeration algorithm with output linear delay of this set of subsequences, with preprocessing time O(w)O(|w|), which is further improved to an incremental enumeration algorithm with O(1)O(1) delay of this set of subsequences, with preprocessing time O(w)O(|w|).

Keywords

Cite

@article{arxiv.2407.18599,
  title  = {Tight Bounds for the Number of Absent Subsequences},
  author = {Duncan Adamson and Pamela Fleischmann and Annika Huch and Florin Manea and Paul Sarnighausen-Cahn and Max Wiedenhöft},
  journal= {arXiv preprint arXiv:2407.18599},
  year   = {2025}
}