English
Related papers

Related papers: Density of rational languages under shift invarian…

200 papers

In this article we undertake a study of extension complexity from the perspective of formal languages. We define a natural way to associate a family of polytopes with binary languages. This allows us to define the notion of extension…

Computational Complexity · Computer Science 2019-08-29 Hans Raj Tiwary

We investigate the number of sets of words that can be formed from a finite alphabet, counted by the total length of the words in the set. An explicit expression for the counting sequence is derived from the generating function, and…

Combinatorics · Mathematics 2010-01-26 Stefan Gerhold

We prove existence of asymptotic entropy of random walks on regular languages over a finite alphabet and we give formulas for it. Furthermore, we show that the entropy varies real-analytically in terms of probability measures of constant…

Probability · Mathematics 2015-03-11 Lorenz A. Gilch

Formal languages are sets of strings of symbols described by a set of rules specific to them. In this note, we discuss a certain class of formal languages, called regular languages, and put forward some elementary results. The properties of…

Formal Languages and Automata Theory · Computer Science 2020-05-22 Aalok Thakkar

Recently, researchers started to pay attention to the detection of temporal shifts in the meaning of words. However, most (if not all) of these approaches restricted their efforts to uncovering change over time, thus neglecting other…

Computation and Language · Computer Science 2017-11-16 Hosein Azarbonyad , Mostafa Dehghani , Kaspar Beelen , Alexandra Arkut , Maarten Marx , Jaap Kamps

The Swadesh approach for determining the temporal separation between two languages relies on the stochastic process of words replacement (when a complete new word emerges to represent a given concept). It is well known that the basic…

Computation and Language · Computer Science 2025-10-28 Maurizio Serva

We study ergodic-theoretic properties of coded shift spaces. A coded shift space is defined as a closure of all bi-infinite concatenations of words from a fixed countable generating set. We derive sufficient conditions for the uniqueness of…

Dynamical Systems · Mathematics 2024-07-11 Tamara Kucherenko , Martin Schmoll , Christian Wolf

The words of a language are randomly replaced in time by new ones, but it has long been known that words corresponding to some items (meanings) are less frequently replaced than others. Usually, the rate of replacement for a given item is…

Computation and Language · Computer Science 2018-10-24 Michele Pasquini , Maurizio Serva

Word embeddings are powerful representations that form the foundation of many natural language processing architectures, both in English and in other languages. To gain further insight into word embeddings, we explore their stability (e.g.,…

Computation and Language · Computer Science 2021-09-13 Laura Burdick , Jonathan K. Kummerfeld , Rada Mihalcea

Some aspects of the physical nature of language are discussed. In particular, physical models of language must exist that are efficiently implementable. The existence requirement is essential because without physical models no communication…

Quantum Physics · Physics 2007-05-23 Paul Benioff

Based on data from a large-scale experiment with human subjects, we conclude that the logarithm of probability to guess a word in context (unpredictability) depends linearly on the word length. This result holds both for poetry and prose,…

Information Theory · Computer Science 2007-07-16 Dmitrii Manin

The concept of "lost positions" is a recently introduced tool for counting the number of runs in words. We investigate the frequency of lost positions in prefixes of words. This leads to an algorithm that allows to show, using an extensive…

Formal Languages and Automata Theory · Computer Science 2019-12-18 Štěpán Holub

In this work we introduce the concepts of linguistic transformation, linguistic loop and semantic deficit. By exploiting Lie group theoretical and geometric techniques, we define invariants that capture the structural properties of a whole…

Computation and Language · Computer Science 2025-04-01 Daniele Corradetti , Alessio Marrani

Given a finite set of words w1,...,wn independently drawn according to a fixed unknown distribution law P called a stochastic language, an usual goal in Grammatical Inference is to infer an estimate of P in some class of probabilistic…

Machine Learning · Computer Science 2007-05-23 François Denis , Yann Esposito , Amaury Habrard

We review recent progress in understanding the meaning of mutual information in natural language. Let us define words in a text as strings that occur sufficiently often. In a few previous papers, we have shown that a power-law distribution…

Information Theory · Computer Science 2020-03-11 Łukasz Dębowski

We study possibilities for semantic and syntactic rigidity, i.e., the rigidity with respect to automorphism group and with respect to definable closure. Variations of rigidity and their degrees are studied in general case, for special…

Logic · Mathematics 2023-07-26 Sergey V. Sudoplatov

We relate two measures of complexity of regular languages. The first is syntactic complexity, that is, the cardinality of the syntactic semigroup of the language. That semigroup is isomorphic to the semigroup of transformations of states…

Formal Languages and Automata Theory · Computer Science 2013-05-24 Janusz Brzozowski , Gareth Davies

Over the last million years, human language has emerged and evolved as a fundamental instrument of social communication and semiotic representation. People use language in part to convey emotional information, leading to the central and…

We deal with finitely additive measures defined on all subsets of natural numbers which extend the asymptotic density (density measures). We consider a class of density measures which are constructed from free ultrafilters on natural…

Number Theory · Mathematics 2016-09-02 Ryoichi Kunisada

The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…

Computation and Language · Computer Science 2012-07-17 Reginald D. Smith