中文
相关论文

相关论文: Distinct word length frequencies: distributions an…

200 篇论文

The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This increase is measured…

计算与语言 · 计算机科学 2017-10-04 Xiaoyong Yan , Petter Minnhagen

The problem addressed concerns the determination of the average number of successive attempts of guessing a word of a certain length consisting of letters with given probabilities of occurrence. Both first- and second-order approximations…

信息论 · 计算机科学 2015-06-19 Kerstin Andersson

We demonstrate that the frequency distribution of phonemes across languages can be explained at both macroscopic and microscopic levels. Macroscopically, phoneme rank-frequency distributions closely follow the order statistics of a…

计算与语言 · 计算机科学 2026-03-04 Fermín Moscoso del Prado Martín , Suchir Salhan

The inverse relationship between the length of a word and the frequency of its use, first identified by G.K. Zipf in 1935, is a classic empirical law that holds across a wide range of human languages. We demonstrate that length is one…

计算与语言 · 计算机科学 2017-06-02 Stephan C. Meylan , Thomas L. Griffiths

Given a random text over a finite alphabet, we study the frequencies at which fixed-length words occur as subsequences. As the data size grows, the joint distribution of word counts exhibits a rich asymptotic structure. We investigate all…

概率论 · 数学 2026-05-06 Chaim Even-Zohar , Tsviqa Lakrec , Ran J. Tessler

The average uncertainty associated with words is an information-theoretic concept at the heart of quantitative and computational linguistics. The entropy has been established as a measure of this average uncertainty - also called average…

计算与语言 · 计算机科学 2016-06-23 Christian Bentz , Dimitrios Alikaniotis

As is the case of many signals produced by complex systems, language presents a statistical structure that is balanced between order and disorder. Here we review and extend recent results from quantitative characterisations of the degree of…

计算与语言 · 计算机科学 2015-03-05 Marcelo A Montemurro , Damián H Zanette

We consider the number of occurrences of subwords (non-consecutive sub-sequences) in a given word. We first define the notion of subword entropy of a given word that measures the maximal number of occurrences among all possible subwords. We…

组合数学 · 数学 2025-10-06 Wenjie Fang

We study the frequency distributions and correlations of the word lengths of ten European languages. Our findings indicate that a) the word-length distribution of short words quantified by the mean value and the entropy distinguishes the…

Recent observations in the theory of verse and empirical metrics have suggested that constructing a verse line involves a pattern-matching search through a source text, and that the number of found elements (complete words totaling a…

cmp-lg · 计算机科学 2007-05-23 Hideaki Aoyama , John Constable

We estimate the $n$-gram entropies of natural language texts in word-length representation and find that these are sensitive to text language and genre. We attribute this sensitivity to changes in the probability distribution of the lengths…

The word-stock of a language is a complex dynamical system in which words can be created, evolve, and become extinct. Even more dynamic are the short-term fluctuations in word usage by individuals in a population. Building on the recent…

物理与社会 · 物理学 2013-04-09 Eduardo G. Altmann , Zakary L. Whichard , Adilson E. Motter

We focus on the statistics of word occurrences and of the waiting times between such occurrences in Blogs. Due to the heterogeneity of words' frequencies, the empirical analysis is performed by studying classes of "frequently-equivalent"…

信息论 · 计算机科学 2012-09-25 R. Lambiotte , M. Ausloos , M. Thelwall

Background: Zipf's discovery that word frequency distributions obey a power law established parallels between biological and physical processes, and language, laying the groundwork for a complex systems perspective on human communication.…

计算与语言 · 计算机科学 2009-11-11 Eduardo G. Altmann , Janet B. Pierrehumbert , Adilson E. Motter

In this article we propose a novel method to estimate the frequency distribution of linguistic variables while controlling for statistical non-independence due to shared ancestry. Unlike previous approaches, our technique uses all available…

种群与进化 · 定量生物学 2021-03-22 Gerhard Jäger , Johannes Wahle

We investigate the number of sets of words that can be formed from a finite alphabet, counted by the total length of the words in the set. An explicit expression for the counting sequence is derived from the generating function, and…

组合数学 · 数学 2010-01-26 Stefan Gerhold

Zipf's law on word frequency is observed in English, French, Spanish, Italian, and so on, yet it does not hold for Chinese, Japanese or Korean characters. A model for writing process is proposed to explain the above difference, which takes…

数据分析、统计与概率 · 物理学 2013-05-03 Linyuan Lu , Zi-Ke Zhang , Tao Zhou

Beyond the local constraints imposed by grammar, words concatenated in long sequences carrying a complex message show statistical regularities that may reflect their linguistic role in the message. In this paper, we perform a systematic…

统计力学 · 物理学 2007-05-23 Marcelo A. Montemurro , Damian H. Zanette

Shannon entropy is often a quantity of interest to linguists studying the communicative capacity of human language. However, entropy must typically be estimated from observed data because researchers do not have access to the underlying…

计算与语言 · 计算机科学 2022-04-06 Aryaman Arora , Clara Meister , Ryan Cotterell

Written language is a complex communication signal capable of conveying information encoded in the form of ordered sequences of words. Beyond the local order ruled by grammar, semantic and thematic structures affect long-range patterns in…

物理与社会 · 物理学 2010-05-17 Marcelo A. Montemurro , Damian Zanette
‹ 上一页 1 2 3 10 下一页 ›