English
Related papers

Related papers: A Statistical Model of Word Rank Evolution

200 papers

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

Statistical studies of languages have focused on the rank-frequency distribution of words. Instead, we introduce here a measure of how word ranks change in time and call this distribution \emph{rank diversity}. We calculate this diversity…

Computation and Language · Computer Science 2015-05-15 Germinal Cocho , Jorge Flores , Carlos Gershenson , Carlos Pineda , Sergio Sánchez

Studies of the overall structure of vocabulary and its dynamics became possible due to creation of diachronic text corpora, especially Google Books Ngram. This article discusses the question of core change rate and the degree to which the…

Computation and Language · Computer Science 2020-03-25 Valery D. Solovyev , Vladimir V. Bochkarev , Anna V. Shevlyakova

The time variation of the rank $k$ of words for six Indo-European languages is obtained using data from Google Books. For low ranks the distinct languages behave differently, maybe due to syntaxis rules, whereas for $k>50$ the law of large…

Physics and Society · Physics 2026-02-04 Germinal Cocho , R. F. Rodríguez , Sergio Sánchez , Jorge Flores , Carlos Pineda , Carlos Gershenson

The word-stock of a language is a complex dynamical system in which words can be created, evolve, and become extinct. Even more dynamic are the short-term fluctuations in word usage by individuals in a population. Building on the recent…

Physics and Society · Physics 2013-04-09 Eduardo G. Altmann , Zakary L. Whichard , Adilson E. Motter

We propose a stochastic model for the number of different words in a given database which incorporates the dependence on the database size and historical changes. The main feature of our model is the existence of two different classes of…

Physics and Society · Physics 2013-05-16 Martin Gerlach , Eduardo G. Altmann

A recent increase in data availability has allowed the possibility to perform different statistical linguistic studies. Here we use the Google Books Ngram dataset to analyze word flow among English, French, German, Italian, and Spanish. We…

Computation and Language · Computer Science 2023-01-18 Josué Ely Molina , Jorge Flores , Carlos Gershenson , Carlos Pineda

We introduce a dataset for studying the evolution of words, constructed from WordNet and the Google Books Ngram Corpus. The dataset tracks the evolution of 4,000 synonym sets (synsets), containing 9,000 English words, from 1800 AD to 2000…

Computation and Language · Computer Science 2019-08-21 Peter D. Turney , Saif M. Mohammad

We show that across architecture (Transformer vs. Mamba vs. RWKV), training dataset (OpenWebText vs. The Pile), and scale (14 million parameters to 12 billion parameters), autoregressive language models exhibit highly consistent patterns of…

Computation and Language · Computer Science 2025-10-30 James A. Michaelov , Roger P. Levy , Benjamin K. Bergen

Understanding how words change their meanings over time is key to models of language and cultural evolution, but historical data on meaning is scarce, making theories hard to develop and test. Word embeddings show promise as a diachronic…

Computation and Language · Computer Science 2018-10-26 William L. Hamilton , Jure Leskovec , Dan Jurafsky

An automatic word classification system has been designed which processes word unigram and bigram frequency statistics extracted from a corpus of natural language utterances. The system implements a binary top-down form of word clustering…

cmp-lg · Computer Science 2016-08-31 John McMahon , F. J. Smith

The words of a language are randomly replaced in time by new ones, but it has long been known that words corresponding to some items (meanings) are less frequently replaced than others. Usually, the rate of replacement for a given item is…

Computation and Language · Computer Science 2018-10-24 Michele Pasquini , Maurizio Serva

Languages vary considerably in syntactic structure. About 40% of the world's languages have subject-verb-object order, and about 40% have subject-object-verb order. Extensive work has sought to explain this word order variation across…

Computation and Language · Computer Science 2022-06-10 Michael Hahn , Yang Xu

We present a probabilistic language model for time-stamped text data which tracks the semantic evolution of individual words over time. The model represents words and contexts by latent trajectories in an embedding space. At each moment in…

Machine Learning · Statistics 2017-07-19 Robert Bamler , Stephan Mandt

Virtually anything can be and is ranked; people, institutions, countries, words, genes. Rankings reduce complex systems to ordered lists, reflecting the ability of their elements to perform relevant functions, and are being used from…

Physics and Society · Physics 2026-02-04 Gerardo Iñiguez , Carlos Pineda , Carlos Gershenson , Albert-László Barabási

Recent neural models have shown significant progress in dialogue generation. Most generation models are based on language models. However, due to the Long Tail Phenomenon in linguistics, the trained models tend to generate words that appear…

Computation and Language · Computer Science 2020-05-05 Zhiqiang Zhan , Zifeng Hou , Yang Zhang

We analyze the dynamic properties of 10^7 words recorded in English, Spanish and Hebrew over the period 1800--2008 in order to gain insight into the coevolution of language and culture. We report language independent patterns useful as…

Physics and Society · Physics 2012-03-16 Alexander M. Petersen , Joel Tenenbaum , Shlomo Havlin , H. Eugene Stanley

Language change is a complex social phenomenon, revealing pathways of communication and sociocultural influence. But, while language change has long been a topic of study in sociolinguistics, traditional linguistic research methods rely on…

Computation and Language · Computer Science 2016-09-08 Rahul Goel , Sandeep Soni , Naman Goyal , John Paparrizos , Hanna Wallach , Fernando Diaz , Jacob Eisenstein

Word embeddings are computed by a class of techniques within natural language processing (NLP), that create continuous vector representations of words in a language from a large text corpus. The stochastic nature of the training process of…

Computation and Language · Computer Science 2020-08-03 Lucas Rettenmeier

Understanding what constitutes high-quality pre-training data remains a central question in language model training. In this work, we investigate whether benchmark performance is primarily driven by the degree of statistical pattern overlap…

Computation and Language · Computer Science 2026-02-12 Woojin Chung , Jeonghoon Kim
‹ Prev 1 2 3 10 Next ›