English
Related papers

Related papers: Dynamics of core of language vocabulary

200 papers

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

The availability of large linguistic data sets enables data-driven approaches to study linguistic change. The Google Books corpus unigram frequency data set is used to investigate the word rank dynamics in eight languages. We observed the…

Computation and Language · Computer Science 2022-02-15 Alex John Quijano , Rick Dale , Suzanne Sindi

Dynamics of average length of words in Russian and English is analysed in the article. Words belonging to the diachronic text corpus Google Books Ngram and dated back to the last two centuries are studied. It was found out that average word…

Computation and Language · Computer Science 2016-05-26 Vladimir V. Bochkarev , Anna V. Shevlyakova , Valery D. Solovyev

We propose a stochastic model for the number of different words in a given database which incorporates the dependence on the database size and historical changes. The main feature of our model is the existence of two different classes of…

Physics and Society · Physics 2013-05-16 Martin Gerlach , Eduardo G. Altmann

Statistical studies of languages have focused on the rank-frequency distribution of words. Instead, we introduce here a measure of how word ranks change in time and call this distribution \emph{rank diversity}. We calculate this diversity…

Computation and Language · Computer Science 2015-05-15 Germinal Cocho , Jorge Flores , Carlos Gershenson , Carlos Pineda , Sergio Sánchez

In this paper, a method for measuring synchronic corpus (dis-)similarity put forward by Kilgarriff (2001) is adapted and extended to identify trends and correlated changes in diachronic text data, using the Corpus of Historical American…

Computation and Language · Computer Science 2015-08-28 Alexander Koplenig

We introduce a dataset for studying the evolution of words, constructed from WordNet and the Google Books Ngram Corpus. The dataset tracks the evolution of 4,000 synonym sets (synsets), containing 9,000 English words, from 1800 AD to 2000…

Computation and Language · Computer Science 2019-08-21 Peter D. Turney , Saif M. Mohammad

It is tempting to treat frequency trends from the Google Books data sets as indicators of the "true" popularity of various words and phrases. Doing so allows us to draw quantitatively strong conclusions about the evolution of cultural…

Physics and Society · Physics 2020-05-28 Eitan Adam Pechenick , Christopher M. Danforth , Peter Sheridan Dodds

All living languages change over time. The causes for this are many, one being the emergence and borrowing of new linguistic elements. Competition between the new elements and older ones with a similar semantic or grammatical function may…

Computation and Language · Computer Science 2020-06-17 Andres Karjus , Richard A. Blythe , Simon Kirby , Kenny Smith

Terms in diachronic text corpora may exhibit a high degree of semantic dynamics that is only partially captured by the common notion of semantic change. The new measure of context volatility that we propose models the degree by which terms…

Computation and Language · Computer Science 2017-11-16 Christian Kahmann , Andreas Niekler , Gerhard Heyer

Diachronic word embeddings -- vector representations of words over time -- offer remarkable insights into the evolution of language and provide a tool for quantifying sociocultural change from text documents. Prior work has used such…

Computation and Language · Computer Science 2020-10-05 Sandeep Soni , Kristina Lerman , Jacob Eisenstein

The words of a language are randomly replaced in time by new ones, but it has long been known that words corresponding to some items (meanings) are less frequently replaced than others. Usually, the rate of replacement for a given item is…

Computation and Language · Computer Science 2018-10-24 Michele Pasquini , Maurizio Serva

This article is devoted to the verification of the empirical Heaps law in European languages using Google Books Ngram corpus data. The connection between word distribution frequency and expected dependence of individual word number on text…

Computation and Language · Computer Science 2020-03-30 Vladimir V. Bochkarev , Eduard Yu. Lerner , Anna V. Shevlyakova

Our languages are in constant flux driven by external factors such as cultural, societal and technological changes, as well as by only partially understood internal motivations. Words acquire new meanings and lose old senses, new words are…

Computation and Language · Computer Science 2019-03-14 Nina Tahmasebi , Lars Borin , Adam Jatowt

A recent increase in data availability has allowed the possibility to perform different statistical linguistic studies. Here we use the Google Books Ngram dataset to analyze word flow among English, French, German, Italian, and Spanish. We…

Computation and Language · Computer Science 2023-01-18 Josué Ely Molina , Jorge Flores , Carlos Gershenson , Carlos Pineda

We present work in progress on the temporal progression of compositionality in noun-noun compounds. Previous work has proposed computational methods for determining the compositionality of compounds. These methods try to automatically…

Computation and Language · Computer Science 2019-06-13 Prajit Dhar , Janis Pagel , Lonneke van der Plas

Word embeddings are powerful representations that form the foundation of many natural language processing architectures, both in English and in other languages. To gain further insight into word embeddings, we explore their stability (e.g.,…

Computation and Language · Computer Science 2021-09-13 Laura Burdick , Jonathan K. Kummerfeld , Rada Mihalcea

Of basic interest is the quantification of the long term growth of a language's lexicon as it develops to more completely cover both a culture's communication requirements and knowledge space. Here, we explore the usage dynamics of words in…

Computation and Language · Computer Science 2017-03-27 Eitan Adam Pechenick , Christopher M. Danforth , Peter Sheridan Dodds

This paper reviews the state-of-the-art of semantic change computation, one emerging research field in computational linguistics, proposing a framework that summarizes the literature by identifying and expounding five essential components…

Computation and Language · Computer Science 2018-06-20 Xuri Tang

The evolution of natural languages poses a riddle to any theoretical perspective based on efficiency considerations. If languages are already optimally effective means of organization and communication of thought, why do they change? And if…

Neurons and Cognition · Quantitative Biology 2026-03-20 Hediye Yarahmadi , Kwang Il Ryom , Giuseppe Longobardi , Alessandro Treves
‹ Prev 1 2 3 10 Next ›