English
Related papers

Related papers: Entropy and Long range correlations in literary En…

200 papers

We analyze the structure of DNA molecules of different organisms by using the additive Markov chain approach. Transforming nucleotide sequences into binary strings, we perform statistical analysis of the corresponding "texts". We develop…

Other Quantitative Biology · Quantitative Biology 2014-11-14 S. S. Melnik , O. V. Usatenko

In this paper we analyse the fractal structure of long human-language records by mapping large samples of texts onto time series. The particular mapping set up in this work is inspired on linguistic basis in the sense that is retains {\em…

Statistical Mechanics · Physics 2007-05-23 Marcelo A. Montemurro , Pedro A. Pury

Given a sequence composed of a limit number of characters, we try to "read" it as a "text". This involves to segment the sequence into "words". The difficulty is to distinguish good segmentation from enormous number of random ones.Aiming at…

Biological Physics · Physics 2009-11-06 Bin Wang

A new approach to estimate the Shannon entropy of a long-range correlated sequence is proposed. The entropy is written as the sum of two terms corresponding respectively to power-law (\emph{ordered}) and exponentially (\emph{disordered})…

Genomics · Quantitative Biology 2013-10-30 Anna Carbone

The frequency with which the letters of the English alphabet appear in writings has been applied to the field of cryptography, the development of keyboard mechanics, and the study of linguistics. We expanded on the statistical analysis of…

Information Theory · Computer Science 2024-01-30 Neil Zhao , Diana Zheng

Hidden structural patterns in written texts have been subject of considerable research in the last decades. In particular, mapping a text into a time series of sentence lengths is a natural way to investigate text structure. Typically,…

Computation and Language · Computer Science 2018-05-07 Denner S. Vieira , Sergio Picoli , Renio S. Mendes

We discuss algorithms for estimating the Shannon entropy h of finite symbol sequences with long range correlations. In particular, we consider algorithms which estimate h from the code lengths produced by some compression algorithm. Our…

Statistical Mechanics · Physics 2017-04-24 Thomas Schürmann , Peter Grassberger

We investigate the longest common substring problem for encoded sequences and its asymptotic behaviour. The main result is a strong law of large numbers for a re-scaled version of this quantity, which presents an explicit relation with the…

Probability · Mathematics 2019-12-12 Adriana Coutinho , Rodrigo Lambert , Jérôme Rousseau

The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This increase is measured…

Computation and Language · Computer Science 2017-10-04 Xiaoyong Yan , Petter Minnhagen

Intermolecular correlations lower values of both diffusion and entropy. We present an analysis of the existing relations between long-time diffusion (D) and entropy. S. A recently proposed inequality, a lower bound, by Sorkin et al.,…

Statistical Mechanics · Physics 2024-06-14 Subhajit Acharya , Biman Bagchi

We present evidence that the word entropy of American English has been rising steadily since around 1900, contrary to predictions from existing sociolinguistic theories. We also find differences in word entropy between media categories,…

General Economics · Economics 2023-04-20 Charlie Pilgrim , Weisi Guo , Thomas T. Hills

This paper investigates the information encoded in the embeddings of large language models (LLMs). We conduct simulations to analyze the representation entropy and discover a power law relationship with model sizes. Building upon this…

Machine Learning · Computer Science 2024-02-07 Zhiquan Tan , Chenghai Li , Weiran Huang

As first shown by H. S. Green in 1952, the entropy of a classical fluid of identical particles can be written as a sum of many-particle contributions, each of them being a distinctive functional of all spatial distribution functions up to a…

Statistical Mechanics · Physics 2021-03-18 Santi Prestipino , Paolo V. Giaquinta

We show that the predictability of letters in written English texts depends strongly on their position in the word. The first letters are usually the least easy to predict. This agrees with the intuitive notion that words are well defined…

Physics and Society · Physics 2017-04-24 Thomas Schürmann , Peter Grassberger

The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…

Computation and Language · Computer Science 2012-07-17 Reginald D. Smith

We study the entropy $S$ of longest increasing subsequences (LIS), i.e., the logarithm of the number of distinct LIS. We consider two ensembles of sequences, namely random permutations of integers and sequences drawn i.i.d.\ from a limited…

Disordered Systems and Neural Networks · Physics 2020-06-09 Phil Krabbe , Hendrik Schawe , Alexander K. Hartmann

Rooted trees with probabilities are used to analyze properties of a variable length code. A bound is derived on the difference between the entropy rates of the code and a memoryless source. The bound is in terms of normalized informational…

Information Theory · Computer Science 2013-10-11 Georg Böcherer , Rana Ali Amjad

The dependence with text length of the statistical properties of word occurrences has long been considered a severe limitation quantitative linguistics. We propose a simple scaling form for the distribution of absolute word frequencies…

Physics and Society · Physics 2015-06-15 Francesc Font-Clos , Gemma Boleda , Álvaro Corral

The origin of long-range letter correlations in natural texts is studied using random walk analysis and Jensen-Shannon divergence. It is concluded that they result from slow variations in letter frequency distribution, which are a…

Computation and Language · Computer Science 2016-11-27 Dmitrii Y. Manin

We study the problem of entropy calibration, which asks whether a language model's entropy over generations matches its log loss on human text. Past work found that models are miscalibrated, with entropy per step increasing as generations…

Computation and Language · Computer Science 2026-01-14 Steven Cao , Gregory Valiant , Percy Liang