Related papers: Finding trends and statistical patterns in name me…
We focus on the statistics of word occurrences and of the waiting times between such occurrences in Blogs. Due to the heterogeneity of words' frequencies, the empirical analysis is performed by studying classes of "frequently-equivalent"…
Complex systems comprise a large number of interacting elements, whose dynamics is not always a priori known. In these cases -- in order to uncover their key features -- we have to turn to empirical methods, one of which was recently…
In this paper we try to model certain features of human language complexity by means of advanced concepts borrowed from statistical mechanics. We use a time series approach, the diffusion entropy method (DE), to compute the complexity of an…
Taylor's law quantifies the scaling properties of the fluctuations of the number of innovations occurring in open systems. Urn based modelling schemes have already proven to be effective in modelling this complex behaviour. Here, we present…
We study the rank distribution, the cumulative probability, and the probability density of returns of stock prices of listed firms traded in four stock markets. We find that the rank distribution and the cumulative probability of stock…
The word-frequency distribution provides the fundamental building blocks that generate discourse in language. It is well known, from empirical evidence, that the word-frequency distribution of almost any text is described by Zipf's law, at…
We study the relationship between vocabulary size and text length in a corpus of $75$ literary works in English, authored by six writers, distinguishing between the contributions of three grammatical classes (or ``tags,'' namely, {\it…
In probabilistic approaches to classification and information extraction, one typically builds a statistical model of words under the assumption that future data will exhibit the same regularities as the training data. In many data sets,…
We empirically study the activity patterns of individual blog-posting and find significant memory effects. The memory coefficient first decays in a power law and then turns to an exponential form. Moreover, the inter-event time distribution…
The dataset was collected to examine and identify possible key topics within these texts. Data preparation such as data cleaning, transformation, tokenization, removal of stop words from both English and Filipino, and word stemming was…
A generic communication model of a boolean network with transmission errors is proposed to explore the power-law scaling of states' evolution in small-world networks. In the model, the power spectrum of the population difference between…
We show that power-law analyses of financial commentaries from newspaper web-sites can be used to identify stock market bubbles, supplementing traditional volatility analyses. Using a four-year corpus of 17,713 online, finance-related…
The temporal statistics exhibited by written correspondence appear to be media dependent, with features which have so far proven difficult to characterize. We explain the origin of these difficulties by disentangling the role of spontaneous…
In this paper, we apply a method to quantify biases associated with named entities from various countries. We create counterfactual examples with small perturbations on target-domain data instead of relying on templates or specific datasets…
In social tagging systems, the diversity of tag vocabulary and the popularity of such tags continue to increase as they are exposed to selection pressure derived from our cognitive nature and cultural preferences. This is analogous to…
The distribution of the number of academic publications as a function of citation count for a given year is remarkably similar from year to year. We measure this similarity as a width of the distribution and find it to be approximately…
We analyze a large-scale snapshot of del.icio.us and investigate how the number of different tags in the system grows as a function of a suitably defined notion of time. We study the temporal evolution of the global vocabulary size, i.e.…
Statistical studies of languages have focused on the rank-frequency distribution of words. Instead, we introduce here a measure of how word ranks change in time and call this distribution \emph{rank diversity}. We calculate this diversity…
We show that the Zipf's law for Chinese characters perfectly holds for sufficiently short texts (few thousand different characters). The scenario of its validity is similar to the Zipf's law for words in short English texts. For long…
We introduce a stochastic model to explain a double power-law distribution which exhibits two different Paretian behaviors in the upper and the lower tail and widely exists in social and economic systems. The model incorporates fitness…