English
Related papers

Related papers: Random Text, Zipf's Law, Critical Length,and Impli…

200 papers

The Random Language Model, proposed as a simple model of human languages, is defined by the averaged model of a probabilistic context-free grammar. This grammar expresses the process of sentence generation as a tree graph with nodes having…

Disordered Systems and Neural Networks · Physics 2022-07-07 Kai Nakaishi , Koji Hukushima

Apparently random events in nature often reveal hidden patterns when analysed using diverse and robust statistical tools. Power-law distributions, for example, project diverse natural phenomenon, ranging from earthquakes1 to heartbeat…

A plethora of natural and socio-economic phenomena share a striking statistical regularity, that is the magnitude of elements decreases with a power law as a function of their position in a ranking of magnitude. Such regularity is known as…

Physics and Society · Physics 2025-10-16 Davide Cugini , André Timpanaro , Giacomo Livan , Giacomo Guarnieri

A topological argument is presented concering the structure of semantic space, based on the negative correlation between polysemy and word length. The resulting graph structure is applied to the modeling of free-recall experiments,…

Neurons and Cognition · Quantitative Biology 2016-11-16 Francesco Fumarola

Syntactic structures used to play a vital role in natural language processing (NLP), but since the deep learning revolution, NLP has been gradually dominated by neural models that do not consider syntactic structures in their design. One…

Computation and Language · Computer Science 2023-11-28 Haoyi Wu , Kewei Tu

We show that the exponent in the inverse power law of word frequencies for the monkey-at-the-typewriter model of Zipf's law will tend towards -1 under broad conditions as the alphabet size increases to infinity and the letter probabilities…

Probability · Mathematics 2015-12-08 Richard Perline

Consider a random word $X^n=(X_1,\ldots ,X_n)$ in an alphabet consisting of $4$ letters, with the letters viewed either as $A$, $U$, $G$ and $C$ (i.e., nucleotides in an RNA sequence) or $\alpha$, $\bar{\alpha}$, $\beta$ and $\bar{\beta}$…

Group Theory · Mathematics 2022-01-20 Siddhartha Gadgil , Manjunath Krishnapur

Tries are among the most versatile and widely used data structures on words. They are pertinent to the (internal) structure of (stored) words and several splitting procedures used in diverse contexts ranging from document taxonomy to IP…

Probability · Mathematics 2012-09-20 Kevin Leckey , Ralph Neininger , Wojciech Szpankowski

The probabilistic Latent Semantic Indexing model assumes that the expectation of the corpus matrix is low-rank and can be written as the product of a topic-word matrix and a word-document matrix. In this paper, we study the estimation of…

Methodology · Statistics 2023-10-11 Huy Tran , Yating Liu , Claire Donnat

Free association is a task that requires a subject to express the first word to come to their mind when presented with a certain cue. It is a task which can be used to expose the basic mechanisms by which humans connect memories. In this…

Adaptation and Self-Organizing Systems · Physics 2014-04-15 Guillermo A. Luduena , M. Djalali Behzad , Claudius Gros

The random Fibonacci chain is a generalisation of the classical Fibonacci substitution and is defined as the rule mapping $0\mapsto 1$ and $1 \mapsto 01$ with probability $p$ and $1 \mapsto 10$ with probability $1-p$ for $0<p<1$ and where…

Combinatorics · Mathematics 2010-01-21 Johan Nilsson

The Random Language Model (De Giuli 2019) is an ensemble of stochastic context-free grammars, quantifying the syntax of human and computer languages. The model suggests a simple picture of first language learning as a type of annealing in…

Disordered Systems and Neural Networks · Physics 2024-12-10 Fatemeh Lalegani , Eric De Giuli

Quantifying the similarity between symbolic sequences is a traditional problem in Information Theory which requires comparing the frequencies of symbols in different sequences. In numerous modern applications, ranging from DNA over music to…

Physics and Society · Physics 2016-04-18 Martin Gerlach , Francesc Font-Clos , Eduardo G. Altmann

A statistical language model assigns probability to strings of arbitrary length. Unfortunately, it is not possible to gather reliable statistics on strings of arbitrary length from a finite corpus. Therefore, a statistical language model…

cmp-lg · Computer Science 2008-02-03 Eric Sven Ristad , Robert G. Thomas

Zipf's law is a paradigm describing the importance of different elements in communication systems, especially in linguistics. Despite the complexity of the hierarchical structure of language, music has in some sense an even more complex…

Physics and Society · Physics 2023-11-20 Marc Serra-Peralta , Joan Serrà , Álvaro Corral

Zipf's law for cities is probably the most famous regularity in social sciences. So much that, a hundred years of publication later, its status is not clear: is it a law of social organisation? Is it an instrument of description of city…

Physics and Society · Physics 2017-10-10 Clementine Cottineau

When following a sequence - such as reading a text or tracking a user's activity - one can measure how the "dictionary" of distinct elements (types) grows with the number of observations (tokens). When this growth follows a power law, it is…

Physics and Society · Physics 2026-04-21 Célestin Zimmerlin , Thomas Louail , Manuel Moussallam , Marc Barthelemy

This paper introduces a statistical and other analysis of peer reviewers in order to approach their "quality" through some quantification measure, thereby leading to some quality metrics. Peer reviewer reports for the Journal of the Serbian…

Physics and Society · Physics 2016-06-08 Marcel Ausloos , Olgica Nedic , Agata Fronczak , Piotr Fronczak

We define and completely solve a content-based directed network whose nodes consist of random words and an adjacency rule involving perfect or approximate matches, for an alphabet with an arbitrary number of letters. The analytic expression…

Molecular Networks · Quantitative Biology 2007-05-23 Muhittin Mungan , Alkan Kabakcioglu , Duygu Balcan , Ayse Erzan

The frequency of the preferred order for a noun phrase formed by demonstrative, numeral, adjective and noun has received significant attention over the last two decades. We investigate the actual distribution of the 24 possible orders.…

Computation and Language · Computer Science 2026-01-23 Ramon Ferrer-i-Cancho
‹ Prev 1 8 9 10 Next ›