相关论文: Autocorrelations Decay in Texts and Applicability …
It is known for some time that autocorrelations of words in human-written texts decay according to a power law. Recent works have also shown that the autocorrelations decay in texts generated by LLMs is qualitatively different from the…
A theory of additive Markov chains with long-range memory is used for description of correlation properties of coarse-grained literary texts. The complex structure of the correlations in texts is revealed. Antipersistent correlations at…
We are interested in investigating the statistical properties of extreme values for strongly correlated variables. The starting motivation is to understand how the strong-correlation properties of power-law distributed processes affect the…
We study synthetic temporal networks whose evolution is determined by stochastically evolving node variables - synthetic analogues of, e.g., temporal proximity networks of mobile agents. We quantify the long-timescale correlations of these…
Automatic extraction of cause-effect relationships from natural language texts is a challenging open problem in Artificial Intelligence. Most of the early attempts at its solution used manually constructed linguistic and syntactic rules on…
Understanding texts requires memory: the reader has to keep in mind enough words to create meaning. This calls for a relation between the memory of the reader and the structure of the text. To investigate this interaction, we first identify…
We show that the mutual information between two symbols, as a function of the number of symbols between the two, decays exponentially in any probabilistic regular grammar, but can decay like a power law for a context-free grammar. This…
Language models must capture statistical dependencies between words at timescales ranging from very short to very long. Earlier work has demonstrated that dependencies in natural language tend to decay with distance between words according…
We propose a power-law decay model with autocorrelation for posting data to social networking services concerning particular events such as national holidays or major sport events. In these kinds of events we observe people's interest both…
This paper presents the results of a study on the semantic constraints imposed on lexical choice by certain contextual indicators. We show how such indicators are computed and how correlations between them and the choice of a noun phrase…
The hydrodynamic part of the velocity autocorrelation function of a granular fluid in the homogeneous cooling state has been calculated by using mode-coupling theory for a finite system with periodic boundary conditions. The existence of…
Current large-scale auto-regressive language models display impressive fluency and can generate convincing text. In this work we start by asking the question: Can the generations of these models be reliably distinguished from real text by…
This paper introduces new methods based on exponential families for modeling the correlations between words in text and speech. While previous work assumed the effects of word co-occurrence statistics to be constant over a window of several…
We consider two graph models of semantic change. The first is a time-series model that relates embedding vectors from one time period to embedding vectors of previous time periods. In the second, we construct one graph for each word: nodes…
The angular and frequency correlation functions of the transmission coefficient for light propagation through a strongly scattering amplifying medium are considered. It is found that just as in the case of an elastic scattering medium the…
Statistical properties of the taxonomic classification of human languages are studied. It is shown that, at the highest levels of the taxonomic hierarchy, the frequency of taxon members as a function of the number of languages belonging to…
Natural languages are full of rules and exceptions. One of the most famous quantitative rules is Zipf's law which states that the frequency of occurrence of a word is approximately inversely proportional to its rank. Though this `law' of…
Large language models (LLMs) have achieved remarkable progress in natural language generation, yet they continue to display puzzling behaviors -- such as repetition and incoherence -- even when exhibiting low perplexity. This highlights a…
Long-term temporal correlations observed in event sequences of natural and social phenomena have been characterized by algebraically decaying autocorrelation functions. Such temporal correlations can be understood not only by heterogeneous…
In this paper we analyse the fractal structure of long human-language records by mapping large samples of texts onto time series. The particular mapping set up in this work is inspired on linguistic basis in the sense that is retains {\em…