English

Languages of Words of Low Automatic Complexity Are Hard to Compute

Formal Languages and Automata Theory 2025-10-10 v1 Logic

Abstract

The automatic complexity of a finite word (string) is an analogue for finite automata of Sipser's distinguishing complexity (1983) and was introduced by Shallit and Wang (2001). For a finite alphabet Σ\Sigma of at least two elements, we consider the non-deterministic automatic complexity given by exactly - yet not necessarily uniquely - accepting automata: a word xΣx \in \Sigma^* has exact non-deterministic automatic complexity kNk \in \mathbb{N} if there exists a non-deterministic automaton of kk states which accepts xx while rejecting every other word of the same length as xx, and no automaton of fewer states has this property. Importantly, and in contrast to the classical notion, the witnessing automaton may have multiple paths of computation accepting xx. We denote this measure of complexity by ANeA_{Ne}, and study a class of languages of low ANeA_{Ne}-complexity defined as Lq={xΣ:ANe(x)<qx}L_q = \{ \, x \in \Sigma^* : A_{Ne}(x) < q|x| \, \}, which is parameterised by rationals q(0,1/2)q \in (0,1/2) (generalising a class of sets first studied by Kjos-Hanssen). We show that for every q(0,1/2)q \in (0,1/2), this class is neither context-free nor recognisable by certain Boolean circuits. In the process, we answer an open question of Kjos-Hanssen quantifying the complexity of L1/3L_{1/3} in terms of Boolean circuits, and also prove the Shannon effect for ANeA_{Ne}.

Keywords

Cite

@article{arxiv.2510.07696,
  title  = {Languages of Words of Low Automatic Complexity Are Hard to Compute},
  author = {Joey Chen and Bjørn Kjos-Hanssen and Ivan Koswara and Linus Richter and Frank Stephan},
  journal= {arXiv preprint arXiv:2510.07696},
  year   = {2025}
}

Comments

22 pages, 1 figure