English
Related papers

Related papers: Entropy and Long range correlations in literary En…

200 papers

Entities involve important concepts with concrete meanings and play important roles in numerous linguistic tasks. Entities have different forms in different linguistic tasks and researchers treat those different forms as different concepts.…

Computation and Language · Computer Science 2023-08-31 Xiaoshi Zhong , Xiang Yu , Erik Cambria , Jagath C. Rajapakse

We propose a compression-based version of the empirical entropy of a finite string over a finite alphabet. Whereas previously one considers the naked entropy of (possibly higher order) Markov processes, we consider the sum of the…

Information Theory · Computer Science 2011-04-05 Paul M. B. Vitányi

Intertextuality is a central tenet in literary studies. It refers to the intricate links between literary texts that are created by various types of references. This paper proposes a new quantitative model of intertextuality to enable…

Computation and Language · Computer Science 2025-09-10 Yi Xing

Whereas for strings, higher-order empirical entropy is the standard entropy measure, several different notions of empirical entropy for trees have been proposed in the past, notably label entropy, degree entropy, conditional versions of the…

Information Theory · Computer Science 2020-06-03 Danny Hucke , Markus Lohrey , Louisa Seelbach Benkner

We describe general approach to classification of character sequences (texts, DNA) using relative entropy estimated by off-the-shelf compression and Markov Chains and find them precise enough. We also notice that the method for estimating…

Statistical Mechanics · Physics 2007-05-23 Dmitry V. Khmelev , William J. Teahan

The information entropies in coordinate and momentum spaces and their sum ($S_r$, $S_k$, $S$) are evaluated for many nuclei using "experimental" densities or/and momentum distributions. The results are compared with the harmonic oscillator…

Nuclear Theory · Physics 2009-11-11 S. E. Massen , V. P. Psonis , A. N. Antonov

Evidence is given for a systematic text-length dependence of the power-law index gamma of a single book. The estimated gamma values are consistent with a monotonic decrease from 2 to 1 with increasing length of a text. A direct connection…

Physics and Society · Physics 2009-12-10 Sebastian Bernhardsson , Luis Enrique Correa da Rocha , Petter Minnhagen

We compared entropy for texts written in natural languages (English, Spanish) and artificial languages (computer software) based on a simple expression for the entropy as a function of message length and specific word diversity. Code text…

Computation and Language · Computer Science 2015-12-03 Gerardo Febres , Klaus Jaffe , Carlos Gershenson

We review the recent progress in the investigation of powerfree words, with particular emphasis on binary cubefree and ternary squarefree words. Besides various bounds on the entropy, we provide bounds on letter frequencies and consider…

Combinatorics · Mathematics 2008-11-14 Uwe Grimm , Manuela Heuer

The word-frequency distribution of a text written by an author is well accounted for by a maximum entropy distribution, the RGF (random group formation)-prediction. The RGF-distribution is completely determined by the a priori values of the…

Physics and Society · Physics 2017-10-03 Xiao-Yong Yan , Petter Minnhagen

It has been shown in a recent publication that words in human-produced English language tend to have an information content close to the conditional entropy. In this paper, we show that the same is true for events in human-produced…

Sound · Computer Science 2022-11-24 Mathias Rose Bjare , Stefan Lattner

Hilberg (1990) supposed that finite-order excess entropy of a random human text is proportional to the square root of the text length. Assuming that Hilberg's hypothesis is true, we derive Guiraud's law, which states that the number of word…

Computation and Language · Computer Science 2020-03-11 Łukasz Dȩbowski

This study investigates the potential influence of Herman Melville reading on his own writings through computational semantic similarity analysis. Using documented records of books known to have been owned or read by Melville, we compare…

Computation and Language · Computer Science 2026-03-17 Nudrat Habib , Elisa Barney Smith , Steven Olsen Smith

We present in this paper a numerical investigation of literary texts by various well-known English writers, covering the first half of the twentieth century, based upon the results obtained through corpus analysis of the texts. A fractal…

Other Condensed Matter · Physics 2009-11-11 L. L. Goncalves , L. B. Goncalves

In natural language using short sentences is considered efficient for communication. However, a text composed exclusively of such sentences looks technical and reads boring. A text composed of long ones, on the other hand, demands…

We develop a theory of classical complexity. We study the relations between classical complexity and entropy, and conjecture that in an isolated system, classical absolute complexity always tends to grow, until it reaches its maximum. We…

High Energy Physics - Theory · Physics 2019-02-28 Zhou Shangnan

Recently, it has been claimed that a linear relationship between a measure of information content and word length is expected from word length optimization and it has been shown that this linearity is supported by a strong correlation…

Data Analysis, Statistics and Probability · Physics 2019-12-11 Ramon Ferrer-i-Cancho , Fermín Moscoso del Prado Martín

We study analytically the corrections to the leading terms in the Renyi entropy of a massive lattice theory, showing significant deviations from naive expectations. In particular, we show that finite size and finite mass effects give rise…

High Energy Physics - Theory · Physics 2012-05-31 Elisa Ercolessi , Stefano Evangelisti , Fabio Franchini , Francesco Ravanini

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

Quantifying the similarity between symbolic sequences is a traditional problem in Information Theory which requires comparing the frequencies of symbols in different sequences. In numerous modern applications, ranging from DNA over music to…

Physics and Society · Physics 2016-04-18 Martin Gerlach , Francesc Font-Clos , Eduardo G. Altmann