Related papers: Randomness versus specifics for word-frequency dis…
We analyse a generalisation of the Galam model of binary opinion dynamics in which iterative discussions take place in local groups of individuals and study the effects of random deviations from the group majority. The probability of a…
We study universal traits which emerge both in real-world complex datasets, as well as in artificially generated ones. Our approach is to analogize data to a physical system and employ tools from statistical physics and Random Matrix Theory…
For words, rank-frequency distributions have long been heralded for adherence to a potentially-universal phenomenon known as Zipf's law. The hypothetical form of this empirical phenomenon was refined by Ben\^{i}ot Mandelbrot to that which…
Words in natural language follow a Zipfian distribution whereby some words are frequent but most are rare. Learning representations for words in the "long tail" of this distribution requires enormous amounts of data. Representations of rare…
The task of finding a criterion allowing to distinguish a text from an arbitrary set of words is rather relevant in itself, for instance, in the aspect of development of means for internet-content indexing or separating signals and noise in…
In this paper, the joint distribution of the sum and maximum of independent, not necessarily identically distributed, nonnegative random variables is studied for two cases: i) continuous and ii) discrete random variables. First, a recursive…
Probability distributions and densities are derived for the excess and deficiency of the intensity or instantaneous energy (quasi-static power) associated with a $p$-dimensional random vector field. Explicit expressions for the exact…
The word-stock of a language is a complex dynamical system in which words can be created, evolve, and become extinct. Even more dynamic are the short-term fluctuations in word usage by individuals in a population. Building on the recent…
We study the girth of Cayley graphs of finite classical groups G on random sets of generators. Our main tool is an essentially best possible bound we obtain on the probability that a given word w takes the value 1 when evaluated in G in…
Zipf's law, which states that the probability of an observation is inversely proportional to its rank, has been observed in many domains. While there are models that explain Zipf's law in each of them, those explanations are typically…
The predictions of text classifiers are often driven by spurious correlations -- e.g., the term `Spielberg' correlates with positively reviewed movies, even though the term itself does not semantically convey a positive sentiment. In this…
The normal distribution is used as a unified probability distribution, however, our researcher found that it is not good agreed with the real-life dynamical system's data. We collected and analyzed representative naturally occurring data…
When generating text from probabilistic models, the chosen decoding strategy has a profound effect on the resulting text. Yet the properties elicited by various decoding strategies do not always transfer across natural language generation…
The different between the inverse power function and the negative exponential function is significant. The former suggests a complex distribution, while the latter indicates a simple distribution. However, the association of the power-law…
Systems as diverse as genetic networks or the world wide web are best described as networks with complex topology. A common property of many large networks is that the vertex connectivities follow a scale-free power-law distribution. This…
The paper revisits the investment simulation based on strategies exhibited by Generalized (m,2)-Zipf law to present an interesting characterization of the wildness in financial time series. The investigations of dominant strategies on each…
Two large, open source software systems are analyzed from the vantage point of complex adaptive systems theory. For both systems, the full dependency graphs are constructed and their properties are shown to be consistent with the assumption…
A theoretical framework is proposed for the understanding of verbal perception -- the conversion of words into meaning, modeled as a compromise between lexical demands and contextual constraints -- and the theory is tested against…
It is the common lore to assume that knowing the equation for the probability distribution function (PDF) of a stochastic model as a function of time tells the whole picture defining all other characteristics of the model. We show that this…
When following a sequence - such as reading a text or tracking a user's activity - one can measure how the "dictionary" of distinct elements (types) grows with the number of observations (tokens). When this growth follows a power law, it is…