Related papers: Entropy and Long range correlations in literary En…
We analyze the rank-frequency distributions of words in selected English and Polish texts. We compare scaling properties of these distributions in both languages. We also study a few small corpora of Polish literary texts and find that for…
We investigate the problem of percolation of words in a random environment. To each vertex, we independently assign a letter $0$ or $1$ according to Bernoulli r.v.'s with parameter $p$. The environment is the resulting graph obtained from…
Constant entropy rate (conditional entropies must remain constant as the sequence length increases) and uniform information density (conditional probabilities must remain constant as the sequence length increases) are two information…
Quantum uncertainty relations are formulated in terms of relative entropy between distributions of measurement outcomes and suitable reference distributions with maximum entropy. This type of entropic uncertainty relation can be applied…
We consider two notions describing how one finite graph may be larger than another. Using them, we prove several theorems for such pairs that compare the number of spanning trees, the return probabilities of random walks, and the number of…
We carry out a comprehensive analysis of letter frequencies in contemporary written Marathi. We determine sets of letters which statistically predominate any large generic Marathi text, and use these sets to estimate the entropy of Marathi.
We propose a novel approach framed in terms of information theory and entropy to tackle the issue of conspiracy theories propagation. We start with the report of an event (such as 9/11 terroristic attack) represented as a series of…
The sequence $(x_n)_{n\in\mathbb N} = (2,5,15,51,187,\dots)$ given by the rule $x_n=(2^n+1)(2^{n-1}+1)/3$ appears in several seemingly unrelated areas of mathematics. For example, $x_n$ is the density of a language of words of length $n$…
We briefly survey some concepts related to empirical entropy -- normal numbers, de Bruijn sequences and Markov processes -- and investigate how well it approximates Kolmogorov complexity. Our results suggest $\ell$th-order empirical entropy…
This study re-evaluates the assumption that long-range correlations in sentence length are a fundamental feature of natural language and a marker of literary style. While previous research has suggested that punctuation marks--particularly…
We enumerate all ternary length-l square-free words, which are words avoiding squares of words up to length l, for l<=24. We analyse the singular behaviour of the corresponding generating functions. This leads to new upper entropy bounds…
Stemming from relationships between a number of information describing a system and entropy content of the system it is possible to determine maximal cosmological time. The contribution manifests a compatibility of the superstring theory…
Written language is a complex communication signal capable of conveying information encoded in the form of ordered sequences of words. Beyond the local order ruled by grammar, semantic and thematic structures affect long-range patterns in…
This paper proposes a method for measuring semantic similarity between words as a new tool for text analysis. The similarity is measured on a semantic network constructed systematically from a subset of the English dictionary, LDOCE…
Information-theoretic measures such as relative entropy and correlation are extremely useful when modeling or analyzing the interaction of probabilistic systems. We survey the quantum generalization of 5 such measures and point out some of…
Entropy relates the fast, microscopic behaviour of the elements in a system to its slow, macroscopic state. We propose to use it to explain how, as complexity theory suggests, small scale decisions of individuals form cities. For this, we…
The use of statistical methods to analyze large databases of text has been useful to unveil patterns of human behavior and establish historical links between cultures and languages. In this study, we identify literary movements by treating…
We relate the information entropy and the mass variance of any distribution in the regime of small fluctuations. We use a set of Monte Carlo simulations of different homogeneous and inhomogeneous distributions to verify the relation and…
To date, most investigations on surprisal and entropy effects in reading have been conducted on the group level, disregarding individual differences. In this work, we revisit the predictive power of surprisal and entropy measures estimated…
A statistical physics study of punctuation effects on sentence lengths is presented for written texts: {\it Alice in wonderland} and {\it Through a looking glass}. The translation of the first text into esperanto is also considered as a…