English
Related papers

Related papers: A Crucial Parameter for Rank-Frequency Relation in…

200 papers

Let $f (\cdot)$ be the absolute frequency of words and $r$ be the rank of words in decreasing order of frequency, then the following function can fit the rank-frequency relation \[ f (r;s,t) = \left(\frac{r_{\tt max}}{r}\right)^{1-s}…

Computation and Language · Computer Science 2022-05-03 Chenchen Ding

The Chapter starts with introductory information about quantitative linguistics notions, like rank--frequency dependence, Zipf's law, frequency spectra, etc. Similarities in distributions of words in texts with level occupation in quantum…

Data Analysis, Statistics and Probability · Physics 2024-01-04 Andrij Rovenchak

A sequence of real numbers $\{x_{n}\}_{n\in \mathbb{N}}$ is said to be $\alpha \beta$-statistically convergent of order $\gamma$ (where $0<\gamma\leq 1$) to a real number $x$ \cite{a} if for every $\delta>0,$ $$\underset{n\rightarrow…

Probability · Mathematics 2016-05-23 Pratulananda Das , Sanjoy Ghosal , Vatan Karakaya , Sumit Som

Language Models (LMs) have shown promising performance in natural language generation. However, as LMs often generate incorrect or hallucinated responses, it is crucial to correctly quantify their uncertainty in responding to given inputs.…

Computation and Language · Computer Science 2024-09-17 Xinmeng Huang , Shuo Li , Mengxin Yu , Matteo Sesia , Hamed Hassani , Insup Lee , Osbert Bastani , Edgar Dobriban

Zipf's law predicts a power-law relationship between word rank and frequency in language communication systems and has been widely reported in a variety of natural language processing applications. However, the emergence of natural language…

Computation and Language · Computer Science 2018-12-05 Bohdan Khomtchouk , Shyam Sudhakaran

We propose a stochastic model for the number of different words in a given database which incorporates the dependence on the database size and historical changes. The main feature of our model is the existence of two different classes of…

Physics and Society · Physics 2013-05-16 Martin Gerlach , Eduardo G. Altmann

Language identification is a critical component of language processing pipelines (Jauhiainen et al.,2019) and is not a solved problem in real-world settings. We present a lightweight and effective language identifier that is robust to…

Computation and Language · Computer Science 2021-09-22 Dominic Widdows , Chris Brew

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

We study rank-frequency relations for phonemes, the minimal units that still relate to linguistic meaning. We show that these relations can be described by the Dirichlet distribution, a direct analogue of the ideal-gas model in statistical…

Computation and Language · Computer Science 2016-04-22 Weibing Deng , Armen E. Allahverdyan

This paper presents the Nataf-Beta Random Field Classifier, a discriminative approach that extends the applicability of the Beta conjugate prior to classification problems. The approach's key feature is to model the probability of a class…

Machine Learning · Computer Science 2015-04-20 James-A. Goulet

Given an input sequence (or prefix), modern language models often assign high probabilities to output sequences that are repetitive, incoherent, or irrelevant to the prefix; as such, model-generated text also contains such artifacts. To…

Computation and Language · Computer Science 2022-11-16 Kalpesh Krishna , Yapei Chang , John Wieting , Mohit Iyyer

We present the new empirical parameter $f_c$, the most probable usage frequency of a word in a language, computed via the distribution of documents over frequency $x$ of the word. This parameter allows for filtering the core lexicon of a…

Disordered Systems and Neural Networks · Physics 2007-05-23 Dmitri Volchenkov , Philippe Blanchard , Serge Sharoff

In this work we derive the covariant and gauge invariant perturbation equations in general theories of $f(R)$ gravity in the Palatini formalism to linear order and calculate the cosmic microwave background (CMB) and matter power spectra for…

Astrophysics · Physics 2008-11-26 Baojiu Li , K. C. Chan , M. -C. Chu

The Random Language Model, proposed as a simple model of human languages, is defined by the averaged model of a probabilistic context-free grammar. This grammar expresses the process of sentence generation as a tree graph with nodes having…

Disordered Systems and Neural Networks · Physics 2022-07-07 Kai Nakaishi , Koji Hukushima

Recent studies show that Generative Relevance Feedback (GRF), using text generated by Large Language Models (LLMs), can enhance the effectiveness of query expansion. However, LLMs can generate irrelevant information that harms retrieval…

Information Retrieval · Computer Science 2023-06-19 Iain Mackie , Ivan Sekulic , Shubham Chatterjee , Jeffrey Dalton , Fabio Crestani

Let ${\bf L}$ be the unit exponential random variable and ${\bf Z}_\alpha$ the standard positive $\alpha$-stable random variable. We prove that $\{(1-\alpha) \alpha^{\gamma_\alpha} {\bf Z}_\alpha^{-\gamma_\alpha}, 0< \alpha <1\}$ is…

Probability · Mathematics 2014-01-28 Thomas Simon

f(R) gravity is thought to be an alternative to dark energy which can explain the acceleration of the universe. It has been tested by different observations including type Ia supernovae (SNIa), the cosmic microwave background (CMB), the…

Cosmology and Nongalactic Astrophysics · Physics 2012-07-12 Kai Liao , Zong-Hong Zhu

Why does Zipf's law give a good description of data from seemingly completely unrelated phenomena? Here it is argued that the reason is that they can all be described as outcomes of a ubiquitous random group division: the elements can be…

Physics and Society · Physics 2015-03-19 Seung Ki Baek , Sebastian Bernhardsson , Petter Minnhagen

We study a deliberately simple, fully non-linguistic model of text: a sequence of independent draws from a finite alphabet of letters plus a single space symbol. A word is defined as a maximal block of non-space symbols. Within this…

Computation and Language · Computer Science 2025-11-25 Vladimir Berman

The availability of large linguistic data sets enables data-driven approaches to study linguistic change. The Google Books corpus unigram frequency data set is used to investigate the word rank dynamics in eight languages. We observed the…

Computation and Language · Computer Science 2022-02-15 Alex John Quijano , Rick Dale , Suzanne Sindi
‹ Prev 1 2 3 10 Next ›