English
Related papers

Related papers: The word entropy of natural languages

200 papers

Classification is a machine learning method used in many practical applications: text mining, handwritten character recognition, face recognition, pattern classification, scene labeling, computer vision, natural langage processing. A…

Machine Learning · Computer Science 2025-11-05 Doulaye Dembélé

In this paper we quantify the consistency of word usage in written texts represented by complex networks, where words were taken as nodes, by measuring the degree of preservation of the node neighborhood.} Words were considered highly…

Physics and Society · Physics 2013-02-19 Diego R. Amancio , Osvaldo N. Oliveira , Luciano da F. Costa

Words of estimative probability (WEPs), such as ''maybe'' or ''probably not'' are ubiquitous in natural language for communicating estimative uncertainty, compared with direct statements involving numerical probability. Human estimative…

Computation and Language · Computer Science 2024-05-27 Zhisheng Tang , Ke Shen , Mayank Kejriwal

Entropy and differential entropy are important quantities in information theory. A tractable extension to singular random variables-which are neither discrete nor continuous-has not been available so far. Here, we present such an extension…

Information Theory · Computer Science 2017-01-04 Günther Koliander , Georg Pichler , Erwin Riegler , Franz Hlawatsch

We present methods for calculating a measure of phonotactic complexity---bits per phoneme---that permits a straightforward cross-linguistic comparison. When given a word, represented as a sequence of phonemic segments such as symbols in the…

Computation and Language · Computer Science 2020-05-11 Tiago Pimentel , Brian Roark , Ryan Cotterell

Data stream mining problem has caused widely concerns in the area of machine learning and data mining. In some recent studies, ensemble classification has been widely used in concept drift detection, however, most of them regard…

Data Structures and Algorithms · Computer Science 2017-08-14 Junhong Wang , Shuliang Xu , Bingqian Duan , Caifeng Liu , Jiye Liang

We propose a new interpretation of measures of information and disorder by connecting these concepts to group theory in a new way. Entropy and group theory are connected here by their common relation to sets of permutations. A combinatorial…

Information Theory · Computer Science 2019-11-25 David J. Galas

Probability theory is fundamental for modeling uncertainty, with traditional probabilities being real and non-negative. Complex probability extends this concept by allowing complex-valued probabilities, opening new avenues for analysis in…

Information Theory · Computer Science 2025-03-07 Chan Li , Hejun Xu , Zhu Cao

The concept of information has emerged as a language in its own right, bridging several disciplines that analyze natural phenomena and man-made systems. Integrated information has been introduced as a metric to quantify the amount of…

Neurons and Cognition · Quantitative Biology 2019-06-10 Alberto Hernández-Espinosa , Héctor Zenil , Narsis A. Kiani , Jesper Tegnér

A general notion of information-related complexity applicable to both natural and man-made systems is proposed. The overall approach is to explicitly consider a rational agent performing a certain task with a quantifiable degree of success.…

Data Analysis, Statistics and Probability · Physics 2013-01-18 Eugene Perevalov , David Grace

We present two novel models of document coherence and their application to information retrieval (IR). Both models approximate document coherence using discourse entities, e.g. the subject or object of a sentence. Our first model views text…

Information Retrieval · Computer Science 2016-10-31 Casper Petersen , Christina Lioma , Jakob Grue Simonsen , Birger Larsen

We advocate the use of a notion of entropy that reflects the relative abundances of the symbols in an alphabet, as well as the similarities between them. This concept was originally introduced in theoretical ecology to study the diversity…

Machine Learning · Computer Science 2022-07-12 Jose Gallego-Posada , Ankit Vani , Max Schwarzer , Simon Lacoste-Julien

The paper analyzes the entropy of a system composed by non-interacting and indistinguishable particles whose quantum state numbers are modelled as independent and identically distributed classical random variables. The crucial observation…

Statistical Mechanics · Physics 2023-05-18 Arnaldo Spalvieri

We measured entropy and symbolic diversity for English and Spanish texts including literature Nobel laureates and other famous authors. Entropy, symbol diversity and symbol frequency profiles were compared for these four groups. We also…

Computation and Language · Computer Science 2017-01-17 Gerardo Febres , Klaus Jaffe

Well-established automatic analyses of texts mainly consider frequencies of linguistic units, e.g. letters, words and bigrams, while methods based on co-occurrence networks consider the structure of texts regardless of the nodes label (i.e.…

Computation and Language · Computer Science 2018-02-27 Camilo Akimushkin , Diego R. Amancio , Osvaldo N. Oliveira

We use large language models (LLMs) to uncover long-ranged structure in English texts from a variety of sources. The conditional entropy or code length in many cases continues to decrease with context length at least to $N\sim 10^4$…

Statistical Mechanics · Physics 2026-01-01 Colin Scheibner , Lindsay M. Smith , William Bialek

Human communication is commonly represented as a temporal social network, and evaluated in terms of its uniqueness. We propose a set of new entropy-based measures for human communication dynamics represented within the temporal social…

Social and Information Networks · Computer Science 2018-10-29 Marcin Kulisiewicz , Przemysław Kazienko , Bolesław K. Szymański , Radosław Michalski

An approach is proposed to quantify, in bits of information, the actual relevance of analogies in analogy tests. The main component of this approach is a softaccuracy estimator that also yields entropy estimates with compensated biases.…

Computation and Language · Computer Science 2022-10-19 Jugurta Montalvão

Rooted trees with probabilities are used to analyze properties of a variable length code. A bound is derived on the difference between the entropy rates of the code and a memoryless source. The bound is in terms of normalized informational…

Information Theory · Computer Science 2013-10-11 Georg Böcherer , Rana Ali Amjad

The principle of maximum entropy is a broadly applicable technique for computing a distribution with the least amount of information possible constrained to match empirical data, for instance, feature expectations. We seek to generalize…

Information Theory · Computer Science 2022-05-30 Kenneth Bogert