English
Related papers

Related papers: The empirical structure of word frequency distribu…

200 papers

Words in natural language follow a Zipfian distribution whereby some words are frequent but most are rare. Learning representations for words in the "long tail" of this distribution requires enormous amounts of data. Representations of rare…

Machine Learning · Computer Science 2018-03-08 Dzmitry Bahdanau , Tom Bosc , Stanisław Jastrzębski , Edward Grefenstette , Pascal Vincent , Yoshua Bengio

The uniform information density (UID) hypothesis proposes that speakers aim to distribute information evenly throughout a text, balancing production effort and listener comprehension difficulty. However, language typically does not maintain…

Computation and Language · Computer Science 2025-06-05 Eleftheria Tsipidi , Samuel Kiegeland , Franz Nowak , Tianyang Xu , Ethan Wilcox , Alex Warstadt , Ryan Cotterell , Mario Giulianelli

The review summarizes the main methodological concepts used in studying natural language from the perspective of complexity science and documents their applicability in identifying both universal and system-specific features of language in…

Physics and Society · Physics 2024-01-09 Tomasz Stanisz , Stanisław Drożdż , Jarosław Kwapień

Human languages vary widely in how they encode information within circumscribed semantic domains (e.g., time, space, color, human body parts and activities), but little is known about the global structure of semantic information and nothing…

Computation and Language · Computer Science 2024-02-19 Pedro Aceves , James A. Evans

History-dependent processes are ubiquitous in natural and social systems. Many such stochastic processes, especially those that are associated with complex systems, become more constrained as they unfold, meaning that their sample-space, or…

Physics and Society · Physics 2015-04-16 Bernat Corominas-Murtra , Rudolf Hanel , Stefan Thurner

In this paper we derive the maximum entropy characteristics of a particular rank order distribution, namely the discrete generalized beta distribution, which has recently been observed to be extremely useful in modelling many several…

Physics and Society · Physics 2019-09-30 Abhik Ghosh , Preety Shreya , Banasri Basu

The article introduces corrections to Zipf's and Heaps' laws based on systematic models of the proportion of hapaxes, i.e., words that occur once. The derivation rests on two assumptions: The first one is the standard urn model which…

Computation and Language · Computer Science 2025-05-27 Łukasz Dębowski

Sentence formation is a highly structured, history-dependent, and sample-space reducing (SSR) process. While the first word in a sentence can be chosen from the entire vocabulary, typically, the freedom of choosing subsequent words gets…

Computation and Language · Computer Science 2018-12-31 Rudolf Hanel , Stefan Thurner

We present empirical data on frequency and pattern of misprints in citations to twelve high-profile papers. We find that the distribution of misprints, ranked by frequency of their repetition, follows Zipf's law. We propose a stochastic…

Disordered Systems and Neural Networks · Physics 2007-05-23 M. V. Simkin , V. P. Roychowdhury

Scaling properties of language are a useful tool for understanding generative processes in texts. We investigate the scaling relations in citywise Twitter corpora coming from the Metropolitan and Micropolitan Statistical Areas of the United…

Physics and Society · Physics 2019-03-12 Eszter Bokányi , Dániel Kondor , Gábor Vattay

Inverse power-law interaction forms, such as the inverse-square law, recur across a wide range of physical, social, and spatial systems. While traditionally derived from specific microscopic mechanisms, the ubiquity of these laws suggests a…

Statistical Mechanics · Physics 2025-12-16 Jerome Baray

Recent works have highlighted optimization difficulties faced by gradient descent in training the first and last layers of transformer-based language models, which are overcome by optimizers such as Adam. These works suggest that the…

Machine Learning · Computer Science 2025-05-27 Frederik Kunstner , Francis Bach

Symbolic sequences such as written language and genomic DNA display characteristic frequency distributions and long-range correlations extending over many symbols. In language, this takes the form of Zipf's law for word frequencies together…

Computation and Language · Computer Science 2026-03-04 Marcelo A. Montemurro , Mirko Degli Esposti

Phoneme frequency distributions exhibit robust statistical regularities across languages, including exponential-tailed rank-frequency patterns and a negative relationship between phonemic inventory size and the relative entropy of the…

Computation and Language · Computer Science 2026-03-11 Fermín Moscoso del Prado Martín , Suchir Salhan

Keywords in scientific articles have found their significance in information filtering and classification. In this article, we empirically investigated statistical characteristics and evolutionary properties of keywords in a very famous…

Data Analysis, Statistics and Probability · Physics 2009-06-23 Zike Zhang , Linyuan Lv , Jian-Guo Liu , Tao Zhou

Recent observations in the theory of verse and empirical metrics have suggested that constructing a verse line involves a pattern-matching search through a source text, and that the number of found elements (complete words totaling a…

cmp-lg · Computer Science 2007-05-23 Hideaki Aoyama , John Constable

We explore a probabilistic model of an artistic text: words of the text are chosen independently of each other in accordance with a discrete probability distribution on an infinite dictionary. The words are enumerated 1, 2, $\ldots$, and…

Statistics Theory · Mathematics 2019-05-02 Mikhail Chebunin , Artyom Kovalevskii

The distributions of the number of occurrences of words (the distributions of words for short) play key roles in information theory, statistics, probability theory, ergodic theory, computer science, and DNA analysis. Bassino et al. 2010 and…

Information Theory · Computer Science 2022-11-16 Hayato Takahashi

Researchers have observed that the frequencies of leading digits in many man-made and naturally occurring datasets follow a logarithmic curve, with digits that start with the number 1 accounting for $\sim 30\%$ of all numbers in the dataset…

Computation and Language · Computer Science 2022-12-22 Leo Hsu , Visar Berisha

We explore which linguistic factors -- at the sentence and token level -- play an important role in influencing language model predictions, and investigate whether these are reflective of results found in humans and human corpora (Gries and…

Computation and Language · Computer Science 2024-09-18 Jaap Jumelet , Willem Zuidema , Arabella Sinclair
‹ Prev 1 8 9 10 Next ›