Related papers: Identifying Quantum Mechanical Statistics in Itali…
The inverse relationship between the length of a word and the frequency of its use, first identified by G.K. Zipf in 1935, is a classic empirical law that holds across a wide range of human languages. We demonstrate that length is one…
We study the possibility of applying statistical mechanics to generally covariant quantum theories with a vanishing Hamiltonian. We show that (under certain appropiate conditions) this makes sense, in spite of the absence of a notion of…
We consider the classical problem of discrete distribution estimation using i.i.d. samples in a novel scenario where additional side information is available on the distribution. In large alphabet datasets such as text corpora, such side…
Analyzing football score data with statistical techniques, we investigate how the highly co-operative nature of the game is reflected in averaged properties such as the distributions of scored goals for the home and away teams. It turns out…
Combining intuitive probabilistic assumptions with the basic laws of classical thermodynamics, using the latter to express probabilistic parameters in terms of the thermodynamic quantities, we get a simple unified derivation of the…
Newberry et al. (Detecting evolutionary forces in language change, Nature 551, 2017) tackle an important but difficult problem in linguistics, the testing of selective theories of language change against a null model of drift. Having…
The classical and quantum evolution of a generic probability distribution is analyzed. To that end, a formalism based on the decomposition of the distribution in terms of its statistical moments is used, which makes explicit the differences…
In this work, Bernoulli's Law of Large Numbers, also known as the Golden theorem, has been extended to study the relations between empirical probability and empirical randomness of an otherwise random experiment. Using the example of a coin…
In this article we derive a useful expectation identity using the language of quantum statistical mechanics, where density matrices represent the state of knowledge about the system. This identity allows to establish relations between…
The entropy rate of printed English is famously estimated to be about one bit per character, a benchmark that modern large language models (LLMs) have only recently approached. This entropy rate implies that English contains nearly 80…
We present a methodological framework to discover linguistic and discursive patterns associated to different social groups through contrastive synthetic text generation and statistical analysis. In contrast with previous approaches, we aim…
The physics of randomness and regularities for languages (mother tongues) and their lifetimes and family trees and for the second languages are studied in terms of two opposite processes; random multiplicative noise [1], and fragmentation…
In recent decades it was established that the quantum measurements of physical quantities in space-time points divided by space-like intervals may be correlated. Though such correlation follows from the formulas of quantum mechanics its…
Methods for learning word representations using large text corpora have received much attention lately due to their impressive performance in numerous natural language processing (NLP) tasks such as, semantic similarity measurement, and…
Speech is a distinctive complex feature of human capabilities. In order to understand the physics underlying speech production, in this work we empirically analyse the statistics of large human speech datasets ranging several languages. We…
Categorical compositional distributional model of Coecke et al. (2010) suggests a way to combine grammatical composition of the formal, type logical models with the corpus based, empirical word representations of distributional semantics.…
In this work we develop a series of techniques to quantify the presence of bias and censorship in newspapers. These algorithms are tested analyzing the occurrence of keywords `killed' and `suicide' (`morti', `suicidio' in Italian) and their…
Word similarity has many applications to social science and cultural analytics tasks like measuring meaning change over time and making sense of contested terms. Yet traditional similarity methods based on cosine similarity between word…
Large language models (LLMs) achieve impressive results in terms of fluency in text generation, yet the nature of their linguistic knowledge - in particular the human-likeness of their internal lexicon - remains uncertain. This study…
This paper describes experiments showing that some tasks in natural language processing (NLP) can already be performed using quantum computers, though so far only with small datasets. We demonstrate various approaches to topic…