English
Related papers

Related papers: A statistical test for the Zipf's law by deviation…

200 papers

Natural languages exhibit striking regularities in their statistical structure, including notably the emergence of Zipf's and Heaps' laws. Despite this, it remains broadly unclear how these properties relate to the modern tokenisation…

Computation and Language · Computer Science 2026-01-08 David S. Berman , Alexander G. Stapleton

English words and the outputs of many other natural processes are well-known to follow a Zipf distribution. Yet this thoroughly-established property has never been shown to help compress or predict these important processes. We show that…

Information Theory · Computer Science 2015-05-04 Moein Falahatgar , Ashkan Jafarpour , Alon Orlitsky , Venkatadheeraj Pichapati , Ananda Theertha Suresh

From a grammar point of view, the role of punctuation marks in a sentence is formally defined and well understood. In semantic analysis punctuation plays also a crucial role as a method of avoiding ambiguity of the meaning. A different…

Computation and Language · Computer Science 2016-11-03 Andrzej Kulig , Jaroslaw Kwapien , Tomasz Stanisz , Stanislaw Drozdz

When the probability of measuring a particular value of some quantity varies inversely as a power of that value, the quantity is said to follow a power law, also known variously as Zipf's law or the Pareto distribution. Power laws appear…

Statistical Mechanics · Physics 2019-09-23 M. E. J. Newman

There are different ways of measuring diversity in complex systems. In particular, in language, lexical diversity is characterized in terms of the type-token ratio and the word entropy. We here investigate both diversity metrics in six…

Computation and Language · Computer Science 2025-07-16 Pablo Rosillo-Rodes , Maxi San Miguel , David Sanchez

The frequency of the preferred order for a noun phrase formed by demonstrative, numeral, adjective and noun has received significant attention over the last two decades. We investigate the actual distribution of the 24 possible orders.…

Computation and Language · Computer Science 2026-01-23 Ramon Ferrer-i-Cancho

We show that the laws of Zipf and Benford, obeyed by scores of numerical data generated by many and diverse kinds of natural phenomena and human activity are related to the focal expression of a generalized thermodynamic structure. This…

Statistical Mechanics · Physics 2015-05-19 Carlo Altamirano , Alberto Robledo

In this article, I conduct a textual and contextual analysis of the empirical literature on Zipf's law for cities. Building on previous meta-analysis material openly available, I collect full texts and bibliographies of 66 scientific…

Physics and Society · Physics 2022-02-01 Clémentine Cottineau

In this paper we combine statistical analysis of large text databases and simple stochastic models to explain the appearance of scaling laws in the statistics of word frequencies. Besides the sublinear scaling of the vocabulary size with…

Physics and Society · Physics 2014-11-05 Martin Gerlach , Eduardo G. Altmann

Present human languages display slightly asymmetric log-normal (Gauss) distribution for size [1-3], whereas present cities follow power law (Pareto-Zipf law)[4]. Our model considers the competition between languages and that between cities…

Computational Physics · Physics 2007-05-23 Caglar Tuncay

Conversation is a cornerstone of social connection and is linked to well-being outcomes. Conversations vary widely in type with some portion generating complex, dynamic stories. One approach to studying how conversations unfold in time is…

As is the case of many signals produced by complex systems, language presents a statistical structure that is balanced between order and disorder. Here we review and extend recent results from quantitative characterisations of the degree of…

Computation and Language · Computer Science 2015-03-05 Marcelo A Montemurro , Damián H Zanette

We propose a new method for the calculation of the statistical properties, as e.g. the entropy, of unknown generators of symbolic sequences. The probability distribution $p(k)$ of the elements $k$ of a population can be approximated by the…

chao-dyn · Physics 2009-10-28 Thorsten Pöschel , Werner Ebeling , Helge Rosé

The performance of deep learning in natural language processing has been spectacular, but the reasons for this success remain unclear because of the inherent complexity of deep learning. This paper provides empirical evidence of its…

Computation and Language · Computer Science 2018-02-07 Shuntaro Takahashi , Kumiko Tanaka-Ishii

We consider the problem of inferring the probability distribution associated with a language, given data consisting of an infinite sequence of elements of the languge. We do this under two assumptions on the algorithms concerned: (i) like a…

Machine Learning · Computer Science 2014-07-16 Paul M. B. Vitanyi , Nick Chater

A curious observation was made that the rank statistics of scientific citation numbers follows Zipf-Mandelbrot's law. The same pow-like behavior is exhibited by some simple random citation models. The observed regularity indicates not so…

Physics and Society · Physics 2007-05-23 Z. K. Silagadze

The analysis of strings of $n$ random variables with geometric distribution has recently attracted renewed interest: Archibald et al. consider the number of distinct adjacent pairs in geometrically distributed words. They obtain the…

Probability · Mathematics 2024-02-14 Guy Louchard , Werner Schachinger , Mark Daniel Ward

Heap's Law states that in a large enough text corpus, the number of types as a function of tokens grows as $N=KM^\beta$ for some free parameters $K,\beta$. Much has been written about how this result and various generalizations can be…

Computation and Language · Computer Science 2019-01-04 Victor Davis

Cooking is a cultural expression of human creativity that transcends geography and time through the orchestration of ingredients and techniques, much like languages do through words and syntax. Yet, beneath the apparent diversity of…

The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…

Computation and Language · Computer Science 2012-07-17 Reginald D. Smith
‹ Prev 1 4 5 6 7 8 10 Next ›