English
Related papers

Related papers: Verifying Heaps' law using Google Books Ngram data

200 papers

Zipf's law is a hallmark of several complex systems with a modular structure, such as books composed by words or genomes composed by genes. In these component systems, Zipf's law describes the empirical power law distribution of component…

Statistical Mechanics · Physics 2018-12-05 Andrea Mazzolini , Alberto Colliva , Michele Caselle , Matteo Osella

We inspect the deductive connection between the neural scaling law and Zipf's law -- two statements discussed in machine learning and quantitative linguistics. The neural scaling law describes how the cross entropy rate of a foundation…

Information Theory · Computer Science 2025-12-23 Łukasz Dębowski

We analyze the occurrence frequencies of over 15 million words recorded in millions of books published during the past two centuries in seven different languages. For all languages and chronological subsets of the data we confirm that two…

Physics and Society · Physics 2012-12-12 Alexander M. Petersen , Joel N. Tenenbaum , Shlomo Havlin , H. Eugene Stanley , Matjaz Perc

We study a deliberately simple, fully non-linguistic model of text: a sequence of independent draws from a finite alphabet of letters plus a single space symbol. A word is defined as a maximal block of non-space symbols. Within this…

Computation and Language · Computer Science 2025-11-25 Vladimir Berman

A "monkey book" is a book consisting of a random distribution of letters and blanks, where a group of letters surrounded by two blanks is defined as a word. We compare the statistics of the word distribution for a monkey book with the…

Data Analysis, Statistics and Probability · Physics 2011-09-09 Sebastian Bernhardsson , Seung Ki Baek , Petter Minnhagen

The World Wide Web has grown so big, in such an anarchic fashion, that it is difficult to describe. One of the evident intrinsic characteristics of the World Wide Web is its multilinguality. Here, we present a technique for estimating the…

Computation and Language · Computer Science 2021-08-23 Gregory Grefenstette , Julien Nioche

Recently long range correlations were detected in nucleotide sequences and in human writings by several authors. We undertake here a systematic investigation of two books, Moby Dick by H. Melville and Grimm's tales, with respect to the…

chao-dyn · Physics 2009-10-22 Werner Ebeling , Thorsten Pöschel

Quantifying the similarity between symbolic sequences is a traditional problem in Information Theory which requires comparing the frequencies of symbols in different sequences. In numerous modern applications, ranging from DNA over music to…

Physics and Society · Physics 2016-04-18 Martin Gerlach , Francesc Font-Clos , Eduardo G. Altmann

We focus on the statistics of word occurrences and of the waiting times between such occurrences in Blogs. Due to the heterogeneity of words' frequencies, the empirical analysis is performed by studying classes of "frequently-equivalent"…

Information Theory · Computer Science 2012-09-25 R. Lambiotte , M. Ausloos , M. Thelwall

In this article, I conduct a textual and contextual analysis of the empirical literature on Zipf's law for cities. Building on previous meta-analysis material openly available, I collect full texts and bibliographies of 66 scientific…

Physics and Society · Physics 2022-02-01 Clémentine Cottineau

We provide a method for automatically detecting change in language across time through a chronologically trained neural language model. We train the model on the Google Books Ngram corpus to obtain word vector representations specific to…

Computation and Language · Computer Science 2014-08-26 Yoon Kim , Yi-I Chiu , Kentaro Hanaki , Darshan Hegde , Slav Petrov

To uncover underlying mechanism of collective human dynamics, we survey more than 1.8 billion blog entries and observe the statistical properties of word appearances. We focus on words that show dynamic growth and decay with a tendency to…

Physics and Society · Physics 2015-03-19 Yukie Sano , Kenta Yamada , Hayafumi Watanabe , Hideki Takayasu , Misako Takayasu

Beyond the local constraints imposed by grammar, words concatenated in long sequences carrying a complex message show statistical regularities that may reflect their linguistic role in the message. In this paper, we perform a systematic…

Statistical Mechanics · Physics 2007-05-23 Marcelo A. Montemurro , Damian H. Zanette

Scaling properties of language are a useful tool for understanding generative processes in texts. We investigate the scaling relations in citywise Twitter corpora coming from the Metropolitan and Micropolitan Statistical Areas of the United…

Physics and Society · Physics 2019-03-12 Eszter Bokányi , Dániel Kondor , Gábor Vattay

Given a discrete distribution, an interesting problem is to determine the minimum size of a random sample drawn from this distribution, in order to observe a given number of different records. This problem is related with many applied…

Probability · Mathematics 2012-09-21 Marco Ferrante , Nadia Frigo

Natural languages are full of rules and exceptions. One of the most famous quantitative rules is Zipf's law which states that the frequency of occurrence of a word is approximately inversely proportional to its rank. Though this `law' of…

Computation and Language · Computer Science 2015-05-27 Jake Ryland Williams , James P. Bagrow , Christopher M. Danforth , Peter Sheridan Dodds

The task of text segmentation may be undertaken at many levels in text analysis---paragraphs, sentences, words, or even letters. Here, we focus on a relatively fine scale of segmentation, hypothesizing it to be in accord with a stochastic…

In this work we analyze statistical properties of 91 relatively small texts in 7 different languages (Spanish, English, French, German, Turkish, Russian, Icelandic) as well as texts with randomly inserted spaces. Despite the size (around…

Physics and Society · Physics 2020-07-15 Diego Espitia , Hernán Larralde

We propose a new computational approach for tracking and detecting statistically significant linguistic shifts in the meaning and usage of words. Such linguistic shifts are especially prevalent on the Internet, where the rapid exchange of…

Computation and Language · Computer Science 2014-11-13 Vivek Kulkarni , Rami Al-Rfou , Bryan Perozzi , Steven Skiena

We test the hypothesis that the degree of grammaticalization of German prepositions correlates with their corpus-based contextual dispersion measured by word entropy. We find that there is indeed a moderate correlation for entropy, but a…

Computation and Language · Computer Science 2018-04-19 Dominik Schlechtweg , Sabine Schulte im Walde