English
Related papers

Related papers: Entropy and Long range correlations in literary En…

200 papers

In this paper we analyse the fractal structure of long human-language records by mapping large samples of texts onto time series. The particular mapping set up in this work is inspired on linguistic basis in the sense that is retains {\em…

Statistical Mechanics · Physics 2007-05-23 Marcelo A. Montemurro , Pedro A. Pury

We study the problem of entropy calibration, which asks whether a language model's entropy over generations matches its log loss on human text. Past work found that models are miscalibrated, with entropy per step increasing as generations…

Computation and Language · Computer Science 2026-01-14 Steven Cao , Gregory Valiant , Percy Liang

We study the entropy $S$ of longest increasing subsequences (LIS), i.e., the logarithm of the number of distinct LIS. We consider two ensembles of sequences, namely random permutations of integers and sequences drawn i.i.d.\ from a limited…

Disordered Systems and Neural Networks · Physics 2020-06-09 Phil Krabbe , Hendrik Schawe , Alexander K. Hartmann

We propose a compression-based version of the empirical entropy of a finite string over a finite alphabet. Whereas previously one considers the naked entropy of (possibly higher order) Markov processes, we consider the sum of the…

Information Theory · Computer Science 2011-04-05 Paul M. B. Vitányi

In this study, the output of large language models (LLM) is considered an information source generating an unlimited sequence of symbols drawn from a finite alphabet. Given the probabilistic nature of modern LLMs, we assume a probabilistic…

Computation and Language · Computer Science 2026-02-24 Marco Scharringhausen

We investigate the longest common substring problem for encoded sequences and its asymptotic behaviour. The main result is a strong law of large numbers for a re-scaled version of this quantity, which presents an explicit relation with the…

Probability · Mathematics 2019-12-12 Adriana Coutinho , Rodrigo Lambert , Jérôme Rousseau

We show that the laws of autocorrelations decay in texts are closely related to applicability limits of language models. Using distributional semantics we empirically demonstrate that autocorrelations of words in texts decay according to a…

Computation and Language · Computer Science 2023-05-12 Nikolay Mikhaylovskiy , Ilya Churilov

This paper reports on results on the entropy of the Spanish language. They are based on an analysis of natural language for n-word symbols (n = 1 to 18), trigrams, digrams, and characters. The results obtained in this work are based on the…

Computation and Language · Computer Science 2013-01-15 Fabio G. Guerrero

This paper deals with the construction of the multiparticle correlation expansion of relative entropy for lattice systems. Thanks to this analysis we are able to express the statistical distance between two systems as a series built over…

Statistical Mechanics · Physics 2016-07-08 Marco D'Alessandro

Shannon entropy is often a quantity of interest to linguists studying the communicative capacity of human language. However, entropy must typically be estimated from observed data because researchers do not have access to the underlying…

Computation and Language · Computer Science 2022-04-06 Aryaman Arora , Clara Meister , Ryan Cotterell

This article is devoted to the verification of the empirical Heaps law in European languages using Google Books Ngram corpus data. The connection between word distribution frequency and expected dependence of individual word number on text…

Computation and Language · Computer Science 2020-03-30 Vladimir V. Bochkarev , Eduard Yu. Lerner , Anna V. Shevlyakova

A new approach to estimate the Shannon entropy of a long-range correlated sequence is proposed. The entropy is written as the sum of two terms corresponding respectively to power-law (\emph{ordered}) and exponentially (\emph{disordered})…

Genomics · Quantitative Biology 2013-10-30 Anna Carbone

Intertextuality is a central tenet in literary studies. It refers to the intricate links between literary texts that are created by various types of references. This paper proposes a new quantitative model of intertextuality to enable…

Computation and Language · Computer Science 2025-09-10 Yi Xing

As is the case of many signals produced by complex systems, language presents a statistical structure that is balanced between order and disorder. Here we review and extend recent results from quantitative characterisations of the degree of…

Computation and Language · Computer Science 2015-03-05 Marcelo A Montemurro , Damián H Zanette

Entities involve important concepts with concrete meanings and play important roles in numerous linguistic tasks. Entities have different forms in different linguistic tasks and researchers treat those different forms as different concepts.…

Computation and Language · Computer Science 2023-08-31 Xiaoshi Zhong , Xiang Yu , Erik Cambria , Jagath C. Rajapakse

Recommendations based on behavioral data may be faced with ambiguous statistical evidence. We consider the case of association rules, relevant e.g.~for query and product recommendations. For example: Suppose that a customer belongs to…

Databases · Computer Science 2015-01-12 Rasmus Pagh , Morten Stöckel

We consider the question of entropic uncertainty relations for prime power dimensions. In order to improve upon such uncertainty relations for higher dimensional quantum systems, we derive a tight lower bound amount of entropy for multiple…

Quantum Physics · Physics 2011-10-03 Jakob Funder

This paper presents analysis of 30 literary texts written in English by different authors. For each text, there were created time series representing length of sentences in words and analyzed its fractal properties using two methods of…

Data Analysis, Statistics and Probability · Physics 2015-03-09 Iwona Grabska-Gradzińska , Andrzej Kulig , Jarosław Kwapień , Paweł Oświęcimka , Stanisław Drożdż

Written language is complex. A written text can be considered an attempt to convey a meaningful message which ends up being constrained by language rules, context dependence and highly redundant in its use of resources. Despite all these…

Computation and Language · Computer Science 2019-05-20 E. Estevez-Rams , A. Mesa Rodriguez , D. Estevez-Moya

We compared entropy for texts written in natural languages (English, Spanish) and artificial languages (computer software) based on a simple expression for the entropy as a function of message length and specific word diversity. Code text…

Computation and Language · Computer Science 2015-12-03 Gerardo Febres , Klaus Jaffe , Carlos Gershenson