English
Related papers

Related papers: Word forms - not just their lengths- are optimized…

200 papers

We present a study of morphological irregularity. Following recent work, we define an information-theoretic measure of irregularity based on the predictability of forms in a language. Using a neural transduction model, we estimate this…

Computation and Language · Computer Science 2019-06-28 Shijie Wu , Ryan Cotterell , Timothy J. O'Donnell

Why do children learn some words before others? Understanding individual variability across children and also variability across words, may be informative of the learning processes that underlie language learning. We investigated item-based…

Computation and Language · Computer Science 2021-11-23 Andrew Z. Flores , Jessica Montag , Jon Willits

Statistical properties of the taxonomic classification of human languages are studied. It is shown that, at the highest levels of the taxonomic hierarchy, the frequency of taxon members as a function of the number of languages belonging to…

Adaptation and Self-Organizing Systems · Physics 2007-05-23 Damian H. Zanette

Current evaluation metrics for language modeling and generation rely heavily on the accuracy of predicted (or generated) words as compared to a reference ground truth. While important, token-level accuracy only captures one aspect of a…

Computation and Language · Computer Science 2020-10-15 Shiran Dudy , Steven Bedrick

Evidence is given for a systematic text-length dependence of the power-law index gamma of a single book. The estimated gamma values are consistent with a monotonic decrease from 2 to 1 with increasing length of a text. A direct connection…

Physics and Society · Physics 2009-12-10 Sebastian Bernhardsson , Luis Enrique Correa da Rocha , Petter Minnhagen

The average uncertainty associated with words is an information-theoretic concept at the heart of quantitative and computational linguistics. The entropy has been established as a measure of this average uncertainty - also called average…

Computation and Language · Computer Science 2016-06-23 Christian Bentz , Dimitrios Alikaniotis

We demonstrate that large texts, representing human (English, Russian, Ukrainian) and artificial (C++, Java) languages, display quantitative patterns characterized by the Benford-like and Zipf laws. The frequency of a word following the…

Computation and Language · Computer Science 2018-03-13 Evgeny Shulzinger , Irina Legchenkova , Edward Bormashenko

Zipf's law states that sequential frequencies of words in a text correspond to a power function. Its probabilistic model is an infinite urn scheme with asymptotically power distribution. The exponent of this distribution must be estimated.…

Statistics Theory · Mathematics 2017-06-15 Mikhail Chebunin , Artyom Kovalevskii

While there exist scores of natural languages, each with its unique features and idiosyncrasies, they all share a unifying theme: enabling human communication. We may thus reasonably predict that human cognition shapes how these languages…

Computation and Language · Computer Science 2021-10-01 Tiago Pimentel , Clara Meister , Elizabeth Salesky , Simone Teufel , Damián Blasi , Ryan Cotterell

In this study, we investigate whether speech symbols, learned through deep learning, follow Zipf's law, akin to natural language symbols. Zipf's law is an empirical law that delineates the frequency distribution of words, forming…

Computation and Language · Computer Science 2023-09-19 Shinnosuke Takamichi , Hiroki Maeda , Joonyong Park , Daisuke Saito , Hiroshi Saruwatari

We analyse preference inference, through consistency, for general preference languages based on lexicographic models. We identify a property, which we call strong compositionality, that applies for many natural kinds of preference…

Logic in Computer Science · Computer Science 2024-11-01 Nic Wilson , Anne-Marie George

Words follow the law of brevity, i.e. more frequent words tend to be shorter. From a statistical point of view, this qualitative definition of the law states that word length and word frequency are negatively correlated. Here the recent…

Neurons and Cognition · Quantitative Biology 2014-12-03 Stuart Semple , Minna J. Hsu , Govindasamy Agoramoorthy , Ramon Ferrer-i-Cancho

We checked that the distribution of words in text should uniform, which gives Heaps' law as natural result, that is, the number of types of words can be expressed as a power law of the number of tokens within text. We developed a…

Physics and Society · Physics 2025-04-16 Kim Chol-jun

We explore which linguistic factors -- at the sentence and token level -- play an important role in influencing language model predictions, and investigate whether these are reflective of results found in humans and human corpora (Gries and…

Computation and Language · Computer Science 2024-09-18 Jaap Jumelet , Willem Zuidema , Arabella Sinclair

Fix two words over the binary alphabet $\{0,1\}$, and generate iid Bernoulli$(p)$ bits until one of the words occurs in sequence. This setup, commonly known as Penney's ante, was popularized by Conway, who found (in unpublished work) a…

Combinatorics · Mathematics 2024-10-01 Mathew Drexel , Xuanshan Peng , Jacob Richey

Recently, it has been claimed that a linear relationship between a measure of information content and word length is expected from word length optimization and it has been shown that this linearity is supported by a strong correlation…

Data Analysis, Statistics and Probability · Physics 2019-12-11 Ramon Ferrer-i-Cancho , Fermín Moscoso del Prado Martín

Hidden structural patterns in written texts have been subject of considerable research in the last decades. In particular, mapping a text into a time series of sentence lengths is a natural way to investigate text structure. Typically,…

Computation and Language · Computer Science 2018-05-07 Denner S. Vieira , Sergio Picoli , Renio S. Mendes

Zipf's law in its basic incarnation is an empirical probability distribution governing the frequency of usage of words in a language. As Terence Tao recently remarked, it still lacks a convincing and satisfactory mathematical explanation.…

Information Theory · Computer Science 2013-04-30 Yuri I. Manin

Of basic interest is the quantification of the long term growth of a language's lexicon as it develops to more completely cover both a culture's communication requirements and knowledge space. Here, we explore the usage dynamics of words in…

Computation and Language · Computer Science 2017-03-27 Eitan Adam Pechenick , Christopher M. Danforth , Peter Sheridan Dodds

The word-stock of a language is a complex dynamical system in which words can be created, evolve, and become extinct. Even more dynamic are the short-term fluctuations in word usage by individuals in a population. Building on the recent…

Physics and Society · Physics 2013-04-09 Eduardo G. Altmann , Zakary L. Whichard , Adilson E. Motter
‹ Prev 1 3 4 5 6 7 10 Next ›