Related papers: Scaling Laws in Human Language
The origin(s) of the ubiquity of probability distribution functions (PDF) with power law tails is still a matter of fascination and investigation in many scientific fields from linguistic, social, economic, computer sciences to essentially…
Symbolic sequences such as written language and genomic DNA display characteristic frequency distributions and long-range correlations extending over many symbols. In language, this takes the form of Zipf's law for word frequencies together…
In this article, we investigate the properties of phoneme N-grams across half of the world's languages. We investigate if the sizes of three different N-gram distributions of the world's language families obey a power law. Further, the…
The time evolution of Earth with her cities, languages and countries is considered in terms of the multiplicative noise and the fragmentation- processes, where the related families, size distributions, lifetimes, bilinguals, etc. are…
Zipf's law can be used to describe the rank-size distribution of cities in a region. It was seldom employed to research urban internal structure. In this paper, we demonstrate that the space-filling process within a city follows Zipf's law…
In a prime number decomposition of integers in a given set, the occurrence frequencies of prime numbers are shown to satisfy a general forms of Zipf's law.
Methods and insights from statistical physics are finding an increasing variety of applications where one seeks to understand the emergent properties of a complex interacting system. One such area concerns the dynamics of language at a…
Understanding how language model performance varies with scale is critical to benchmark and algorithm development. Scaling laws are one approach to building this understanding, but the requirement of training models across many different…
We present empirical data on frequency and pattern of misprints in citations to twelve high-profile papers. We find that the distribution of misprints, ranked by frequency of their repetition, follows Zipf's law. We propose a stochastic…
We study rank-size distribution of cities in Japan on the basis of data analysis. From the census data after World War II, we find that the rank-size distribution of cities is composed of two parts, each of which has independent power…
Speech is a distinctive complex feature of human capabilities. In order to understand the physics underlying speech production, in this work we empirically analyse the statistics of large human speech datasets ranging several languages. We…
What processes can explain how very large populations are able to converge on the use of a particular word or grammatical construction without global coordination? Answering this question helps to understand why new language constructs…
Recent studies have shown that as Transformer-based language models become larger and are trained on very large amounts of data, the fit of their surprisal estimates to naturalistic human reading times degrades. The current work presents a…
Evaluating whether large language models (LLMs) capture the structure of natural language beyond local fluency remains an open challenge. Existing evaluation methods, largely based on task performance or short-context behavior, provide…
Statistical studies of languages have focused on the rank-frequency distribution of words. Instead, we introduce here a measure of how word ranks change in time and call this distribution \emph{rank diversity}. We calculate this diversity…
Throughout history most young adults have chosen to live where their parents did while a smaller number moved away. This is sufficient, by proof and simulation, to account for the well-known power law distributions of city sizes. The model…
Corpus-based statistical analysis plays a significant role in linguistic research, and ample evidence has shown that different languages exhibit some common laws. Studies have found that letters in some alphabetic writing languages have…
We study the frequency distributions and correlations of the word lengths of ten European languages. Our findings indicate that a) the word-length distribution of short words quantified by the mean value and the entropy distinguishes the…
We present a comparative analysis of text complexity across domains using scale-free metrics. We quantify linguistic complexity via Heaps' exponent $\beta$ (vocabulary growth), Taylor's exponent $\alpha$ (word-frequency fluctuation…
This paper investigates the rank distribution, cumulative probability, and probability density of price returns for the stocks traded in the KSE and the KOSDAQ market. This research demonstrates that the rank distribution is consistent…