English
Related papers

Related papers: Jumbled Scattered Factors

200 papers

A binary word is Sturmian if the occurrences of each letter are balanced, in the sense that in any two factors of the same length, the difference between the number of occurrences of the same letter is at most 1. In digital geometry,…

Discrete Mathematics · Computer Science 2025-11-11 Alessandro De Luca , Gabriele Fici

We consider a new family of factorial languages whose subword complexity grows as $\Theta(n^{\alpha})$, where $\alpha$ is the root of some transcendent equation. Analytical methods and in particular, a corollary of the Wiener-Pitt theorem,…

Combinatorics · Mathematics 2010-12-30 Julien Cassaigne , Anna Frid , Fedor Petrov

We find generating functions the number of strings (words) containing a specified number of occurrences of certain types of order-isomorphic classes of substrings called subword patterns. In particular, we find generating functions for the…

Combinatorics · Mathematics 2007-05-23 A. Burstein , T. Mansour

We prove that, for any pure morphic word $w$, if the frequencies of all letters in $w$ exist, then the frequencies of all factors in $w$ exist as well. This result answers a question of Saari in his doctoral thesis.

Combinatorics · Mathematics 2024-05-30 Shuo Li

Word embeddings allow natural language processing systems to share statistical information across related words. These embeddings are typically based on distributional statistics, making it difficult for them to generalize to rare or unseen…

Computation and Language · Computer Science 2016-09-27 Parminder Bhatia , Robert Guthrie , Jacob Eisenstein

We consider questions related to the structure of infinite words (over an integer alphabet) with bounded additive complexity, i.e., words with the property that the number of distinct sums exhibited by factors of the same length is bounded…

Combinatorics · Mathematics 2012-09-24 Graham Banero

We consider several sequences of random variables whose Fourier-Laplace transforms present the same type of \textit{splitting phenomenon} when suitably rescaled by the Fourier-Laplace transform of a Poisson-distributed random variable…

Probability · Mathematics 2025-07-24 Yacine Barhoumi-Andréani

A non-empty word $w$ is a border of the word $u$ if $\vert w\vert<\vert u\vert$ and $w$ is both a prefix and a suffix of $u$. A word $u$ with the border $w$ is closed if $u$ has exactly two occurrences of $w$. A word $u$ is privileged if…

Discrete Mathematics · Computer Science 2020-01-22 Josef Rukavicka

We find generating functions for the number of words avoiding certain patterns or sets of patterns on at most 2 distinct letters and determine which of them are equally avoided. We also find the exact number of words avoiding certain…

Combinatorics · Mathematics 2007-05-23 Alexander Burstein , Toufik Mansour

In 2005, Rampersad and the second author proved a number of theorems about infinite words x with the property that if w is any sufficiently long finite factor of x, then its reversal w^R is not a factor of x. In this note we revisit these…

Formal Languages and Automata Theory · Computer Science 2019-12-10 Lukas Fleischer , Jeffrey Shallit

We study some properties of the growth rate of $\mathcal{L}(\mathcal{A},\mathcal{F})$, that is, the language of words over the alphabet $\mathcal{A}$ avoiding the set of forbidden factors $\mathcal{F}$. We first provide a sufficient…

Combinatorics · Mathematics 2025-04-09 Vuong Bui , Matthieu Rosenfeld

Consider a regression or some regression-type model for a certain response variable where the linear predictor includes an ordered factor among the explanatory variables. The inclusion of a factor of this type can take place is a few…

Methodology · Statistics 2023-11-27 Adelchi Azzalini

Word complexity is defined in a number of different ways. Psycholinguistic, morphological and lexical proxies are often used. Human ratings are also used. The problem here is that these proxies do not measure complexity directly, and human…

Computation and Language · Computer Science 2024-08-06 Michael Dalvean

The past years have seen a drastic rise in studies devoted to the investigation of colexification patterns in individual languages families in particular and the languages of the world in specific. Specifically computational studies have…

Computation and Language · Computer Science 2023-02-03 Johann-Mattis List

In this paper we explore a new hierarchy of classes of languages and infinite words and its connection with complexity classes. Namely, we say that a language belongs to the class $L_k$ if it is a subset of the catenation of $k$ languages…

Formal Languages and Automata Theory · Computer Science 2014-06-17 J. Cassaigne , A. E. Frid , S. Puzynina , L. Q. Zamboni

In this paper, we study an abelian-type property of infinite words called well distributed occurrences, or WELLDOC for short. An infinite word $w$ on a $d$-ary alphabet has the WELLDOC property if, for each factor $u$ of $w$, positive…

Discrete Mathematics · Computer Science 2026-03-10 Svetlana Puzynina , Vladimir Schavelev

We characterize words which cluster under the Burrows-Wheeler transform as those words $w$ such that $ww$ occurs in a trajectory of an interval exchange transformation, and build examples of clustering words.

Combinatorics · Mathematics 2012-04-09 Sébastien Ferenczi , Luca Q. Zamboni

This paper introduces new methods based on exponential families for modeling the correlations between words in text and speech. While previous work assumed the effects of word co-occurrence statistics to be constant over a window of several…

cmp-lg · Computer Science 2008-02-03 Doug Beeferman , Adam Berger , John Lafferty

Zipf's law is found when the vocabulary of long written texts is ranked according to the frequency of word occurrences, establishing a power-law decay for the frequency vs rank relation. This law is a robust statistical property observed…

Physics and Society · Physics 2020-02-17 Juan Ignacio Perotti , Orlando Vito Billoni

A pseudo-primitive word with respect to an antimorphic involution \theta is a word which cannot be written as a catenation of occurrences of a strictly shorter word t and \theta(t). Properties of pseudo-primitive words are investigated in…

Computational Complexity · Computer Science 2010-02-23 Lila Kari , Benoît Masson , Shinnosuke Seki