Related papers: Clustering and Arnoux-Rauzy words
A word over an ordered alphabet is said to be clustering if identical letters appear adjacently in its Burrows-Wheeler transform. Such words are strictly related to (discrete) interval exchange transformations. We use an extended version of…
We characterize words which cluster under the Burrows-Wheeler transform as those words $w$ such that $ww$ occurs in a trajectory of an interval exchange transformation, and build examples of clustering words.
The paper deals with balances and imbalances in Arnoux-Rauzy words. We provide sufficient conditions for $C$-balancedness, but our results indicate that even a characterization of 2-balanced Arnoux-Rauzy words on a 3-letter alphabet is not…
We study balancedness properties of words given by the Arnoux-Rauzy and Brun multi-dimensional continued fraction algorithms. We show that almost all Brun words on 3 letters and Arnoux-Rauzy words over arbitrary alphabets are finitely…
Perfectly clustering words are one of many possible generalizations of Christoffel words. In this article, we propose a factorization of a perfectly clustering word on a $n$ letters alphabet into a product of $n-1$ palindromes with a letter…
An infinite word generated by a substitution is rigid if all the substitutions which fix this word are powers of a same substitution. Sturmian words as well as characteristic Arnoux-Rauzy words are known to be rigid. In the present paper,…
We prove that episturmian words and Arnoux-Rauzy sequences can be characterized using a local balance property. We also give a new characterization of epistandard words and show that the set of finite words that are not factors of an…
In this paper, we survey the rich theory of infinite episturmian words which generalize to any finite alphabet, in a rather resembling way, the well-known family of Sturmian words on two letters. After recalling definitions and basic…
In the study of infinite words, various notions of balancedness provide quantitative measures for how regularly letters or factors occur, and they find applications in several areas of mathematics and theoretical computer science. In this…
We construct an Arnoux-Rauzy word for which the set of all differences of two abelianized factors is equal to $\mathbb{Z}^3$. In particular, the imbalance of this word is infinite - and its Rauzy fractal is unbounded in all directions of…
We prove that for every integer $n > 0$ and for every alphabet $\Sigma_k$ of size $k \geq 3$, there exists a necklace of length $n$ whose Burrows-Wheeler Transform (BWT) is completely unclustered, i.e., it consists of exactly $n$ runs with…
Recently, a new characterization of Lyndon words that are also perfectly clustering was proposed by Lapointe and Reutenauer (2024). A word over a ternary alphabet {a,b,c} is called perfectly clustering Lyndon if and only if it is the…
We investigate various connections between the clustering for the Burrows-Wheeler transform, a lossless algorithm used in data compression, and languages of interval exchange transformations. We show that a primitive word $u$ clusters for a…
We describe and experimentally evaluate a method for automatically clustering words according to their distribution in particular syntactic contexts. Deterministic annealing is used to find lowest distortion sets of clusters. As the…
The class of (eventually) dendric words generalizes well-known families such as the Arnoux-Rauzy words or the codings of interval exchanges. There are still many open questions about the link between dendricity and morphisms. In this paper,…
The non-repetitive complexity $nr\mathcal{C}_{\bf u}$ and the initial non-repetitive complexity $inr\mathcal{C}_{\bf u}$ are functions which reflect the structure of the infinite word ${\bf u}$ with respect to the repetitions of factors of…
Text clustering holds significant value across various domains due to its ability to identify patterns and group related information. Current approaches which rely heavily on a computed similarity measure between documents are often limited…
A finite word $u$ is called closed if its longest repeated prefix has exactly two occurrences in $u,$ once as a prefix and once as a suffix. We study the function $f_x^c:\mathbb N \rightarrow \mathbb N$ which counts the number of closed…
We address the problem of clustering words (or constructing a thesaurus) based on co-occurrence data, and using the acquired word classes to improve the accuracy of syntactic disambiguation. We view this problem as that of estimating a…
We consider questions related to the structure of infinite words (over an integer alphabet) with bounded additive complexity, i.e., words with the property that the number of distinct sums exhibited by factors of the same length is bounded…