English
Related papers

Related papers: Random Text, Zipf's Law, Critical Length,and Impli…

200 papers

Word feature vectors have been proven to improve many NLP tasks. With recent advances in unsupervised learning of these feature vectors, it became possible to train it with much more data, which also resulted in better quality of learned…

Computation and Language · Computer Science 2022-11-29 Marius Sajgalik , Michal Barla , Maria Bielikova

It is shown that a real novel shares many characteristic features with a null model in which the words are randomly distributed throughout the text. Such a common feature is a certain translational invariance of the text. Another is that…

Computation and Language · Computer Science 2009-10-14 Sebastian Bernhardsson , Luis Enrique Correa da Rocha , Petter Minnhagen

Text generation tasks, including translation, summarization, language models, and etc. see rapid growth during recent years. Despite the remarkable achievements, the repetition problem has been observed in nearly all text generation models…

Computation and Language · Computer Science 2021-03-23 Zihao Fu , Wai Lam , Anthony Man-Cho So , Bei Shi

In the area of pattern avoidability the central role is played by special words called Zimin patterns. The symbols of these patterns are treated as variables and the rank of the pattern is its number of variables. Zimin type of a word $x$…

Discrete Mathematics · Computer Science 2015-04-01 Wojciech Rytter , Arseny M. Shur

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

We consider languages generated by weighted context-free grammars. It is shown that the behaviour of large texts is controlled by saddle-point equations for an appropriate generating function. We then consider ensembles of grammars, in…

Disordered Systems and Neural Networks · Physics 2022-10-03 Eric De Giuli

This paper addresses the uniform random generation of words from a context-free language (over an alphabet of size $k$), while constraining every letter to a targeted frequency of occurrence. Our approach consists in a multidimensional…

Data Structures and Algorithms · Computer Science 2010-12-21 Olivier Bodini , Yann Ponty

A key aim in biology and psychology is to identify fundamental principles underpinning the behavior of animals, including humans. Analyses of human language and the behavior of a range of non-human animal species have provided evidence for…

Neurons and Cognition · Quantitative Biology 2014-12-03 R. Ferrer-i-Cancho , A. Hernández-Fernández , D. Lusseau , G. Agoramoorthy , M. J. Hsu , S. Semple

What processes can explain how very large populations are able to converge on the use of a particular word or grammatical construction without global coordination? Answering this question helps to understand why new language constructs…

Physics and Society · Physics 2007-05-23 A. Baronchelli , M. Felici , E. Caglioti , V. Loreto , L. Steels

Keywords in scientific articles have found their significance in information filtering and classification. In this article, we empirically investigated statistical characteristics and evolutionary properties of keywords in a very famous…

Data Analysis, Statistics and Probability · Physics 2009-06-23 Zike Zhang , Linyuan Lv , Jian-Guo Liu , Tao Zhou

Statistical studies of languages have focused on the rank-frequency distribution of words. Instead, we introduce here a measure of how word ranks change in time and call this distribution \emph{rank diversity}. We calculate this diversity…

Computation and Language · Computer Science 2015-05-15 Germinal Cocho , Jorge Flores , Carlos Gershenson , Carlos Pineda , Sergio Sánchez

The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…

Computation and Language · Computer Science 2012-07-17 Reginald D. Smith

We analyze the occurrence frequencies of over 15 million words recorded in millions of books published during the past two centuries in seven different languages. For all languages and chronological subsets of the data we confirm that two…

Physics and Society · Physics 2012-12-12 Alexander M. Petersen , Joel N. Tenenbaum , Shlomo Havlin , H. Eugene Stanley , Matjaz Perc

In this paper the Zipf-Mandelbrot law is revisited in the context of linguistics. Despite its widespread popularity the Zipf--Mandelbrot law can only describe the statistical behaviour of a rather restricted fraction of the total number of…

Statistical Mechanics · Physics 2009-11-07 Marcelo A. Montemurro

We describe a novel algorithm for random sampling of freely reduced words equal to the identity in a finitely presented group. The algorithm is based on Metropolis Monte Carlo sampling. The algorithm samples from a stretched Boltzmann…

Group Theory · Mathematics 2013-12-23 M. Elder , A. Rechnitzer , E. J. Janse van Rensburg

A power-free language is characterized by the number of symbols used and a limit on how many times a block of symbols can repeat consecutively. For certain values of these parameters, it is known that the number of legal words grows…

Dynamical Systems · Mathematics 2025-07-28 Vaughn Climenhaga

As is the case of many signals produced by complex systems, language presents a statistical structure that is balanced between order and disorder. Here we review and extend recent results from quantitative characterisations of the degree of…

Computation and Language · Computer Science 2015-03-05 Marcelo A Montemurro , Damián H Zanette

This paper studies the effect of linguistic constraints on the large scale organization of language. It describes the properties of linguistic networks built using texts of written language with the words randomized. These properties are…

Computation and Language · Computer Science 2011-02-16 Madhav Krishna , Ahmed Hassan , Yang Liu , Dragomir Radev

Large Language Models (LLMs) are a powerful tool for statistical text analysis, with derived sequences of next-token probability distributions offering a wealth of information. Extracting this signal typically relies on metrics such as…

A family of information theoretic models of communication was introduced more than a decade ago to explain the origins of Zipf's law for word frequencies. The family is a based on a combination of two information theoretic principles:…

Physics and Society · Physics 2020-09-24 Ramon Ferrer-i-Cancho