English
Related papers

Related papers: Random Text, Zipf's Law, Critical Length,and Impli…

200 papers

By determining which were the most common English words and phrases since the beginning of the 16th century, we obtain a unique large-scale view of the evolution of written text. We find that the most common words and phrases in any given…

Physics and Society · Physics 2012-12-10 Matjaz Perc

We show that statistical criticality, i.e. the occurrence of power law frequency distributions, arises in samples that are maximally informative about the underlying generating process. In order to reach this conclusion, we first identify…

Data Analysis, Statistics and Probability · Physics 2019-07-09 Ryan John Cubero , Junghyo Jo , Matteo Marsili , Yasser Roudi , Juyong Song

As a core element of culture, images transform perception into structured representations and undergo evolution similar to natural languages. Given that visual input accounts for 60% of human sensory experience, it is natural to ask whether…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Ping-Rui Tsai , Chi-hsiang Wang , Yu-Cheng Liao , Hong-Yue Huang , Tzay-Ming Hong

Zipf-like distributions characterize a wide set of phenomena in physics, biology, economics and social sciences. In human activities, Zipf-laws describe for example the frequency of words appearance in a text or the purchases types in…

In his pioneering research, G. K. Zipf formulated a couple of statistical laws on the relationship between the frequency of a word with its number of meanings: the law of meaning distribution, relating the frequency of a word and its…

Computation and Language · Computer Science 2022-01-19 Neus Català , Jaume Baixeries , Ramon Ferrer-Cancho , Lluís Padró , Antoni Hernández-Fernández

The frequency distributions of DNA k-mers are shaped by fundamental biological processes and offer a window into genome structure and evolution. Inspired by analogies to natural language, prior studies have attempted to model genomic k-mer…

n-tuple power law widely exists in language, computer program code, DNA and music. After a vast amount of Zipf analyses of n-tuple power law from empirical data, we propose a model to explain the n-tuple power law feature existed in these…

Data Analysis, Statistics and Probability · Physics 2009-08-05 Xiaocong Gan , Dahui Wang , Zhangang Han

The performance of deep learning in natural language processing has been spectacular, but the reasons for this success remain unclear because of the inherent complexity of deep learning. This paper provides empirical evidence of its…

Computation and Language · Computer Science 2018-02-07 Shuntaro Takahashi , Kumiko Tanaka-Ishii

Understanding the complexity of human language requires an appropriate analysis of the statistical distribution of words in texts. We consider the information retrieval problem of detecting and ranking the relevant words of a text by means…

Computation and Language · Computer Science 2008-06-07 Juan P. Herrera , Pedro A. Pury

Scaling properties of language are a useful tool for understanding generative processes in texts. We investigate the scaling relations in citywise Twitter corpora coming from the Metropolitan and Micropolitan Statistical Areas of the United…

Physics and Society · Physics 2019-03-12 Eszter Bokányi , Dániel Kondor , Gábor Vattay

Complex network theory is used to investigate the structure of meaningful concepts in written texts of individual authors. Networks have been constructed after a two phase filtering, where words with less meaning contents are eliminated,…

Data Analysis, Statistics and Probability · Physics 2009-11-11 Silvia M. G. Caldeira , Thierry C. Petit Lobao , R. F. S. Andrade , Alexis Neme , J. G. V. Miranda

A well-known fact in the field of lossless text compression is that high-order entropy is a weak model when the input contains long repetitions. Motivated by this, decades of research have generated myriads of so-called dictionary…

Data Structures and Algorithms · Computer Science 2020-12-17 Dominik Kempa , Nicola Prezza

In this study, the output of large language models (LLM) is considered an information source generating an unlimited sequence of symbols drawn from a finite alphabet. Given the probabilistic nature of modern LLMs, we assume a probabilistic…

Computation and Language · Computer Science 2026-02-24 Marco Scharringhausen

In this work we analyze statistical properties of 91 relatively small texts in 7 different languages (Spanish, English, French, German, Turkish, Russian, Icelandic) as well as texts with randomly inserted spaces. Despite the size (around…

Physics and Society · Physics 2020-07-15 Diego Espitia , Hernán Larralde

Despite the recent successes of deep learning in natural language processing (NLP), there remains widespread usage of and demand for techniques that do not rely on machine learning. The advantage of these techniques is their…

Computation and Language · Computer Science 2020-12-04 Adam Hare , Yu Chen , Yinan Liu , Zhenming Liu , Christopher G. Brinton

It has been shown recently that a specific class of path-dependent stochastic processes, which reduce their sample space as they unfold, lead to exact scaling laws in frequency and rank distributions. Such Sample Space Reducing processes…

Physics and Society · Physics 2017-10-02 Bernat Corominas-Murtra , Rudolf Hanel , Stefan Thurner

We show that the Zipf's law for Chinese characters perfectly holds for sufficiently short texts (few thousand different characters). The scenario of its validity is similar to the Zipf's law for words in short English texts. For long…

Computation and Language · Computer Science 2014-03-10 W. B. Deng , A. E. Allahverdyan , B. Li , Q. A. Wang

For words, rank-frequency distributions have long been heralded for adherence to a potentially-universal phenomenon known as Zipf's law. The hypothetical form of this empirical phenomenon was refined by Ben\^{i}ot Mandelbrot to that which…

Computation and Language · Computer Science 2017-10-24 Jake Ryland Williams , Giovanni C. Santia

A set of data with positive values follows a Pareto distribution if the log-log plot of value versus rank is approximately a straight line. A Pareto distribution satisfies Zipf's law if the log-log plot has a slope of -1. Since many types…

Economics · Quantitative Finance 2020-06-04 Ricardo T. Fernholz , Robert Fernholz

Traditional studies of memory for meaningful narratives focus on specific stories and their semantic structures but do not address common quantitative features of recall across different narratives. We introduce a statistical ensemble of…

Statistical Mechanics · Physics 2025-02-25 Weishun Zhong , Tankut Can , Antonis Georgiou , Ilya Shnayderman , Mikhail Katkov , Misha Tsodyks