Related papers: Randomness versus specifics for word-frequency dis…
Many dynamical processes on real world networks display complex temporal patterns as, for instance, a fat-tailed distribution of inter-events times, leading to heterogeneous waiting times between events. In this work, we focus on…
Zipf's law is found when the vocabulary of long written texts is ranked according to the frequency of word occurrences, establishing a power-law decay for the frequency vs rank relation. This law is a robust statistical property observed…
We propose a continuum model for the degree distribution of directed networks in free and open-source software. The degree distributions of links in both the in-directed and out-directed dependency networks follow Zipf's law for the…
Motivated by numerous questions in random geometry, given a smooth manifold $M$, we approach a systematic study of the differential topology of Gaussian random fields (GRF) $X:M\to \mathbb{R}^k$, that we interpret as random variables with…
Zipf's law is the most common statistical distribution displaying scaling behavior. Cities, populations or firms are just examples of this seemingly universal law. Although many different models have been proposed, no general theoretical…
Zipf's law predicts a power-law relationship between word rank and frequency in language communication systems, and is widely reported in texts yet remains enigmatic as to its origins. Computer simulations have shown that language…
Language models famously improve under a smooth scaling law, but some specific capabilities exhibit sudden breakthroughs in performance. Advocates of "emergence" view these capabilities as unlocked at a specific scale, but others attribute…
In this paper, the classical problem of the probabilistic characterization of a random variable is re-examined. A random variable is usually described by the probability density function (PDF) or by its Fourier transform, namely the…
In random matrix theory (RMT), the Tracy-Widom (TW) distribution describes the behavior of the largest eigenvalue. We consider here two models in which TW undergoes transformations. In the first one disorder is introduced in the Gaussian…
Diffusion-a measure of dynamics, and entropy-a measure of disorder in the system, are found to be intimately correlated in many systems, and the correlation is often strongly non-linear. We explore the origin of this complex dependence by…
Stopwords are words that are not very informative to the content or the meaning of a language text. Most stopwords are function words but can also be common verbs, adjectives and adverbs. In contrast to the well known Zipf's law for…
We have translated fractional Brownian motion (FBM) signals into a text based on two ''letters'', as if the signal fluctuations correspond to a constant stepsize random walk. We have applied the Zipf method to extract the $\zeta '$ exponent…
The problem addressed concerns the determination of the average number of successive attempts of guessing a word of a certain length consisting of letters with given probabilities of occurrence. Both first- and second-order approximations…
Morphologically rich languages accentuate two properties of distributional vector space models: 1) the difficulty of inducing accurate representations for low-frequency word forms; and 2) insensitivity to distinct lexical relations that…
The article introduces corrections to Zipf's and Heaps' laws based on systematic models of the proportion of hapaxes, i.e., words that occur once. The derivation rests on two assumptions: The first one is the standard urn model which…
In many complex systems, for the activity f(i) of the constituents or nodes i, a power-law relationship was discovered between the standard deviation sigma(i) and the average strength of the activity: sigma(i) ~ <f(i)>^alpha; universal…
Much work has been done on feature selection. Existing methods are based on document frequency, such as Chi-Square Statistic, Information Gain etc. However, these methods have two shortcomings: one is that they are not reliable for…
We study the relationship between vocabulary size and text length in a corpus of $75$ literary works in English, authored by six writers, distinguishing between the contributions of three grammatical classes (or ``tags,'' namely, {\it…
In this paper, I introduce a simple method of computing relative word frequencies for authorship attribution and similar stylometric tasks. Rather than computing relative frequencies as the number of occurrences of a given word divided by…
The problem of continuum percolation in dispersions of rods is reformulated in terms of weighted random geometric graphs. Nodes (or sites or vertices) in the graph represent spatial locations occupied by the centers of the rods. The…