English
Related papers

Related papers: In narrative texts punctuation marks obey the same…

200 papers

Zipf's law states that if words of language are ranked in the order of decreasing frequency in texts, the frequency of a word is inversely proportional to its rank. It is very robust as an experimental observation, but to date it escaped…

Computation and Language · Computer Science 2009-01-22 Dmitrii Manin

In this paper we combine statistical analysis of large text databases and simple stochastic models to explain the appearance of scaling laws in the statistics of word frequencies. Besides the sublinear scaling of the vocabulary size with…

Physics and Society · Physics 2014-11-05 Martin Gerlach , Eduardo G. Altmann

Languages across the world exhibit Zipf's law of abbreviation, namely more frequent words tend to be shorter. The generalized version of the law - an inverse relationship between the frequency of a unit and its magnitude - holds also for…

Information Theory · Computer Science 2016-05-05 R. Ferrer-i-Cancho , C. Bentz , C. Seguin

Free association is a task that requires a subject to express the first word to come to their mind when presented with a certain cue. It is a task which can be used to expose the basic mechanisms by which humans connect memories. In this…

Adaptation and Self-Organizing Systems · Physics 2014-04-15 Guillermo A. Luduena , M. Djalali Behzad , Claudius Gros

The importance of statistical patterns of language has been debated over decades. Although Zipf's law is perhaps the most popular case, recently, Menzerath's law has begun to be involved. Menzerath's law manifests in language, music and…

This paper studies the limits of language models' statistical learning in the context of Zipf's law. First, we demonstrate that Zipf-law token distribution emerges irrespective of the chosen tokenization. Second, we show that Zipf…

Computation and Language · Computer Science 2022-11-22 Elizaveta Zhemchuzhina , Nikolai Filippov , Ivan P. Yamshchikov

In this paper we build on earlier observations and theory regarding word length frequency and sequential distribution to develop a mathematical characterization of some of the language features distinguishing isometrically lineated text…

cmp-lg · Computer Science 2007-05-23 Hideaki Aoyama , John Constable

This paper revisits Menzerath's Law, also known as the Menzerath-Altmann Law, which models a relationship between the length of a linguistic construct and the average length of its constituents. Recent findings indicate that simple…

Computation and Language · Computer Science 2025-10-17 Jiří Milička

The article introduces corrections to Zipf's and Heaps' laws based on systematic models of the proportion of hapaxes, i.e., words that occur once. The derivation rests on two assumptions: The first one is the standard urn model which…

Computation and Language · Computer Science 2025-05-27 Łukasz Dębowski

Sentence formation is a highly structured, history-dependent, and sample-space reducing (SSR) process. While the first word in a sentence can be chosen from the entire vocabulary, typically, the freedom of choosing subsequent words gets…

Computation and Language · Computer Science 2018-12-31 Rudolf Hanel , Stefan Thurner

Statistical methods have been widely employed in recent years to grasp many language properties. The application of such techniques have allowed an improvement of several linguistic applications, which encompasses machine translation,…

Computation and Language · Computer Science 2016-02-22 Henrique F. de Arruda , Luciano da F. Costa , Diego R. Amancio

We checked that the distribution of words in text should uniform, which gives Heaps' law as natural result, that is, the number of types of words can be expressed as a power law of the number of tokens within text. We developed a…

Physics and Society · Physics 2025-04-16 Kim Chol-jun

We show that the Zipf's law for Chinese characters perfectly holds for sufficiently short texts (few thousand different characters). The scenario of its validity is similar to the Zipf's law for words in short English texts. For long…

Computation and Language · Computer Science 2014-03-10 W. B. Deng , A. E. Allahverdyan , B. Li , Q. A. Wang

Zipf's law of abbreviation, the tendency of more frequent words to be shorter, is one of the most solid candidates for a linguistic universal, in the sense that it has the potential for being exceptionless or with a number of exceptions…

Computation and Language · Computer Science 2023-10-13 Sonia Petrini , Antoni Casas-i-Muñoz , Jordi Cluet-i-Martinell , Mengxue Wang , Chris Bentz , Ramon Ferrer-i-Cancho

A stereotype is a generalized perception of a specific group of humans. It is often potentially encoded in human language, which is more common in texts on social issues. Previous works simply define a sentence as stereotypical and…

Computation and Language · Computer Science 2024-01-30 Yang Liu

We observe the statistical properties of blogs that are expected to reflect social human interaction. Firstly, we introduce a basic normalization preprocess that enables us to evaluate the genuine word frequency in blogs that are…

Physics and Society · Physics 2010-04-09 Yukie Sano , Misako Takayasu

Punctuation is a strong indicator of syntactic structure, and parsers trained on text with punctuation often rely heavily on this signal. Punctuation is a diversion, however, since human language processing does not rely on punctuation to…

Computation and Language · Computer Science 2018-09-05 Anders Søgaard , Miryam de Lhoneux , Isabelle Augenstein

Sentence is a basic linguistic unit, however, little is known about how information content is distributed across different positions of a sentence. Based on authentic language data of English, the present study calculated the entropy and…

Computation and Language · Computer Science 2016-09-27 Shuiyuan Yu , Jin Cong , Junying Liang , Haitao Liu

According to Zipf's meaning-frequency law, words that are more frequent tend to have more meanings. Here it is shown that a linear dependency between the frequency of a form and its number of meanings is found in a family of models of…

Computation and Language · Computer Science 2016-10-14 Ramon Ferrer-i-Cancho

While utilizing syntactic tools such as parts-of-speech (POS) tagging has helped us understand sentence structures and their distribution across diverse corpora, it is quite complex and poses a challenge in natural language processing…

Computation and Language · Computer Science 2025-12-15 Abhijeet Sahdev
‹ Prev 1 3 4 5 6 7 10 Next ›