English
Related papers

Related papers: Verifying Heaps' law using Google Books Ngram data

200 papers

It is shown that a real novel shares many characteristic features with a null model in which the words are randomly distributed throughout the text. Such a common feature is a certain translational invariance of the text. Another is that…

Computation and Language · Computer Science 2009-10-14 Sebastian Bernhardsson , Luis Enrique Correa da Rocha , Petter Minnhagen

We investigated long range correlations in two literary texts, Moby Dick by H. Melville and Grimm's tales. The analysis is based on the calculation of entropy like quantities as the mutual information for pairs of letters and the entropy,…

Disordered Systems and Neural Networks · Physics 2015-06-24 Werner Ebeling , Thorsten Poeschel

In this paper we quantify the statistical properties and dynamics of the frequency of hashtag use on Twitter. Hashtags are special words used in social media to attract attention and to organize content. Looking at the collection of all…

Physics and Society · Physics 2020-06-04 Hongjia H. Chen , Tristram J. Alexander , Diego F. M. Oliveira , Eduardo G. Altmann

In this paper, we try to explore the evolution of language through case calculations. First, we chose the novels of eleven British writers from 1400 to 2005 and found the corresponding works; Then, we use the natural language processing…

Computation and Language · Computer Science 2018-10-09 Zhu Gao , Yanhui Jiang , Junhui Gao

Many features from texts and languages can now be inferred from statistical analyses using concepts from complex networks and dynamical systems. In this paper we quantify how topological properties of word co-occurrence networks and…

Physics and Society · Physics 2013-02-20 Diego R. Amancio , Eduardo G. Altmann , Osvaldo N. Oliveira , Luciano da F. Costa

Textual geographic information is indispensable and heavily relied upon in practical applications. The absence of clear distribution poses challenges in effectively harnessing geographic information, thereby driving our quest for…

Computation and Language · Computer Science 2023-09-04 Zhenhua Wang , Daiyu Zhang , Ming Ren , Guang Xu

Hedges are widely studied across registers and disciplines, yet research on the translation of hedges in political texts is extremely limited. This contrastive study is dedicated to investigating whether there is a diachronic change in the…

Computation and Language · Computer Science 2023-05-24 Zhaokun Jiang , Ziyin Zhang

Complex natural and technological systems can be considered, on a coarse-grained level, as assemblies of elementary components: for example, genomes as sets of genes, or texts as sets of words. On one hand, the joint occurrence of…

We investigate the origin of Zipf's law for words in written texts by means of a stochastic dynamical model for text generation. The model incorporates both features related to the general structure of languages and memory effects inherent…

Statistical Mechanics · Physics 2007-05-23 Damián H. Zanette , Marcelo A. Montemurro

Most research related to unithood were conducted as part of a larger effort for the determination of termhood. Consequently, novelties are rare in this small sub-field of term extraction. In addition, existing work were mostly empirically…

Artificial Intelligence · Computer Science 2008-10-02 Wilson Wong , Wei Liu , Mohammed Bennamoun

We study rank-frequency relations for phonemes, the minimal units that still relate to linguistic meaning. We show that these relations can be described by the Dirichlet distribution, a direct analogue of the ideal-gas model in statistical…

Computation and Language · Computer Science 2016-04-22 Weibing Deng , Armen E. Allahverdyan

Zipf's law is a fundamental paradigm in the statistics of written and spoken natural language as well as in other communication systems. We raise the question of the elementary units for which Zipf's law should hold in the most natural way,…

Physics and Society · Physics 2015-07-14 Alvaro Corral , Gemma Boleda , Ramon Ferrer-i-Cancho

The surge in digitized text data requires reliable inferential methods on observed textual patterns. This article proposes a novel two-sample text test for comparing similarity between two groups of documents. The hypothesis is whether the…

Machine Learning · Statistics 2025-05-09 Jingbin Xu , Chen Qian , Meimei Liu , Feng Guo

The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This increase is measured…

Computation and Language · Computer Science 2017-10-04 Xiaoyong Yan , Petter Minnhagen

Understanding the innovation process, that is the underlying mechanisms through which novelties emerge, diffuse and trigger further novelties is undoubtedly of fundamental importance in many areas (biology, linguistics, social science and…

Physics and Society · Physics 2021-10-29 Giacomo Aletti , Irene Crimaldi

In this paper we try to model certain features of human language complexity by means of advanced concepts borrowed from statistical mechanics. We use a time series approach, the diffusion entropy method (DE), to compute the complexity of an…

Statistical Mechanics · Physics 2009-09-03 Paolo Allegrini , Paolo Grigolini , Luigi Palatella

We analyze the rank-frequency distributions of words in selected English and Polish texts. We compare scaling properties of these distributions in both languages. We also study a few small corpora of Polish literary texts and find that for…

Computation and Language · Computer Science 2015-11-11 Stanislaw Drozdz , Jaroslaw Kwapien , Adam Orczyk

The word-frequency distribution of a text written by an author is well accounted for by a maximum entropy distribution, the RGF (random group formation)-prediction. The RGF-distribution is completely determined by the a priori values of the…

Physics and Society · Physics 2017-10-03 Xiao-Yong Yan , Petter Minnhagen

We present a probabilistic language model for time-stamped text data which tracks the semantic evolution of individual words over time. The model represents words and contexts by latent trajectories in an embedding space. At each moment in…

Machine Learning · Statistics 2017-07-19 Robert Bamler , Stephan Mandt

Many studies have shown that human languages tend to optimize for lower complexity and increased communication efficiency. Syntactic dependency distance, which measures the linear distance between dependent words, is often considered a key…

Computation and Language · Computer Science 2024-03-29 Yanran Chen , Wei Zhao , Anne Breitbarth , Manuel Stoeckel , Alexander Mehler , Steffen Eger