Related papers: Maximal State Complexity and Generalized de Bruijn…
Motivated by applications in bioinformatics, we consider the word collector problem, i.e. the expected number of calls to a random weighted generator of words of length $n$ before the full collection is obtained. The originality of this…
Finite chase, or alternatively chase termination, is an important condition to ensure the decidability of existential rule languages. In the past few years, a number of rule languages with finite chase have been studied. In this work, we…
A binary word is a map W : N --> {0,1}, and the set of factors of W with length n is F_n(W):={(W(i),W(i+1),...,W(i+n-1)) : i >= 0}. A word is Sturmian if |F_n(W)|=n+1 for every n>0. We show that the sum of the heights (also known as hamming…
Let $w$ be a finite word of length $n$. In this paper, we study the maximum possible number of distinct rational power factors in a finite word. A rational power is a word of the form $u=p^kp'$, where $p$ is a nonempty finite word, $k$ is…
We study the following problem, first introduced by Dekking. Consider an infinite word x over an alphabet {0,1,...,k-1} and a semigroup homomorphism S:{0,1,...,k-1}* -> N. Let L_x denote the set of factors of x. What conditions on S and the…
This paper presents TextComplexityDE, a dataset consisting of 1000 sentences in German language taken from 23 Wikipedia articles in 3 different article-genres to be used for developing text-complexity predictor models and automatic text…
We solve an open problem concerning syntactic complexity: We prove that the cardinality of the syntactic semigroup of a suffix-free language with $n$ left quotients (that is, with state complexity $n$) is at most $(n-1)^{n-2}+n-2$ for $n\ge…
We prove that, for any arbitrary finite alphabet and for the uniform distribution over deterministic and accessible automata with n states, the average complexity of Moore's state minimization algorithm is in O(n log n). Moreover this bound…
We present a language complexity analysis of World of Warcraft (WoW) community texts, which we compare to texts from a general corpus of web English. Results from several complexity types are presented, including lexical diversity, density,…
We study the optimization version of constraint satisfaction problems (Max-CSPs) in the framework of parameterized complexity; the goal is to compute the maximum fraction of constraints that can be satisfied simultaneously. In standard…
A regular language $L$ is union-free if it can be represented by a regular expression without the union operation. A union-free language is deterministic if it can be accepted by a deterministic one-cycle-free-path finite automaton; this is…
Combinatorial properties of maximal repetitions (runs) in formal words are studied. We classify all maximal repetitions in a word as primary and secondary where the set of all primary repetitions determines all the other repetitons in the…
A reconstruction problem of words from scattered factors asks for the minimal information, like multisets of scattered factors of a given length or the number of occurrences of scattered factors from a given set, necessary to uniquely…
In language learning in the limit, the most common type of hypothesis is to give an enumerator for a language. This so-called $W$-index allows for naming arbitrary computably enumerable languages, with the drawback that even the membership…
We study the problem of generating arithmetic math word problems (MWPs) given a math equation that specifies the mathematical computation and a context that specifies the problem scenario. Existing approaches are prone to generating MWPs…
While many languages possess processes of joining two or more words to create compound words, previous studies have been typically limited only to languages with excessively productive compound formation (e.g., German, Dutch) and there is…
We study the membership problem to context-free languages L (CFLs) on probabilistic words, that specify for each position a probability distribution on the letters (assuming independence across positions). Our task is to compute, given a…
Let $w=w(x_1,...,x_n)$ be a word, i.e. an element of the free group $F = \langle x_1,...,x_n \rangle$. The verbal subgroup $w(G)$ of a group $G$ is the subgroup generated by the set $\{ w(x_1,...,x_n) : x_1,...,x_n \in G \}$ of all…
We introduce a new measure on regular languages: their nondeterministic syntactic complexity. It is the least degree of any extension of the `canonical boolean representation' of the syntactic monoid. Equivalently, it is the least number of…
A group word $w$ is said to be strongly concise in a class $\mathscr C$ of profinite groups if, for any group $G$ in $\mathscr C$, either $w$ takes at least continuum values in $G$ or the verbal subgroup $w(G)$ is finite. It is conjectured…