中文
相关论文

相关论文: On the letter frequencies and entropy of written M…

200 篇论文

This paper presents an investigation of the entropy of the Telugu script. Since this script is syllabic, and not alphabetic, the computation of entropy is somewhat complicated.

计算与语言 · 计算机科学 2011-06-30 Venkata Ravinder Paruchuri

The frequency with which the letters of the English alphabet appear in writings has been applied to the field of cryptography, the development of keyboard mechanics, and the study of linguistics. We expanded on the statistical analysis of…

信息论 · 计算机科学 2024-01-30 Neil Zhao , Diana Zheng

We review the recent progress in the investigation of powerfree words, with particular emphasis on binary cubefree and ternary squarefree words. Besides various bounds on the entropy, we provide bounds on letter frequencies and consider…

组合数学 · 数学 2008-11-14 Uwe Grimm , Manuela Heuer

With the intention of bringing uniformity to Bengali text entry research, here we present a new approach for calculating the most popular English text entry evaluation metrics for Bengali. To demonstrate our approach, we conducted a user…

人机交互 · 计算机科学 2017-06-27 Sayan Sarcar , Ahmed Sabbir Arif , Ali Mazalek

The statistical properties of letters frequencies in European literature texts are investigated. The determination of logarithmic dependence of letters sequence for one-language and two-language texts are examined. The pare of languages is…

We study the entropy of Chinese and English texts, based on characters in case of Chinese texts and based on words for both languages. Significant differences are found between the languages and between different personal styles of debating…

计算与语言 · 计算机科学 2017-01-17 R. R. Xie , W. B. Deng , D. J. Wang , L. P. Csernai

The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…

计算与语言 · 计算机科学 2012-07-17 Reginald D. Smith

We estimate the $n$-gram entropies of natural language texts in word-length representation and find that these are sensitive to text language and genre. We attribute this sensitivity to changes in the probability distribution of the lengths…

We show that the predictability of letters in written English texts depends strongly on their position in the word. The first letters are usually the least easy to predict. This agrees with the intuitive notion that words are well defined…

物理与社会 · 物理学 2017-04-24 Thomas Schürmann , Peter Grassberger

Written language is complex. A written text can be considered an attempt to convey a meaningful message which ends up being constrained by language rules, context dependence and highly redundant in its use of resources. Despite all these…

计算与语言 · 计算机科学 2019-05-20 E. Estevez-Rams , A. Mesa Rodriguez , D. Estevez-Moya

The dependence of the frequency distributions due to multiple meanings of words in a text is investigated by deleting letters. By coding the words with fewer letters the number of meanings per coded word increases. This increase is measured…

计算与语言 · 计算机科学 2017-10-04 Xiaoyong Yan , Petter Minnhagen

Beyond the local constraints imposed by grammar, words concatenated in long sequences carrying a complex message show statistical regularities that may reflect their linguistic role in the message. In this paper, we perform a systematic…

统计力学 · 物理学 2007-05-23 Marcelo A. Montemurro , Damian H. Zanette

A simple method for finding the entropy and redundancy of a reasonable long sample of English text by direct computer processing and from first principles according to Shannon theory is presented. As an example, results on the entropy of…

计算与语言 · 计算机科学 2009-11-19 Fabio G. Guerrero

The word-frequency distribution of a text written by an author is well accounted for by a maximum entropy distribution, the RGF (random group formation)-prediction. The RGF-distribution is completely determined by the a priori values of the…

物理与社会 · 物理学 2017-10-03 Xiao-Yong Yan , Petter Minnhagen

We consider words as a network of interacting letters, and approximate the probability distribution of states taken on by this network. Despite the intuition that the rules of English spelling are highly combinatorial (and arbitrary), we…

神经元与认知 · 定量生物学 2025-02-13 Greg J. Stephens , William Bialek

The average uncertainty associated with words is an information-theoretic concept at the heart of quantitative and computational linguistics. The entropy has been established as a measure of this average uncertainty - also called average…

计算与语言 · 计算机科学 2016-06-23 Christian Bentz , Dimitrios Alikaniotis

The entropy rate of printed English is famously estimated to be about one bit per character, a benchmark that modern large language models (LLMs) have only recently approached. This entropy rate implies that English contains nearly 80…

计算与语言 · 计算机科学 2026-02-19 Weishun Zhong , Doron Sivan , Tankut Can , Mikhail Katkov , Misha Tsodyks

We enumerate all ternary length-l square-free words, which are words avoiding squares of words up to length l, for l<=24. We analyse the singular behaviour of the corresponding generating functions. This leads to new upper entropy bounds…

组合数学 · 数学 2007-05-23 Christoph Richard , Uwe Grimm

Entropy has been a common index to quantify the complexity of time series in a variety of fields. Here, we introduce increment entropy to measure the complexity of time series in which each increment is mapped into a word of two letters,…

数据分析、统计与概率 · 物理学 2016-01-20 Xiaofeng Liu , Aimin Jiang , Ning Xu , Jianru Xue

We compared entropy for texts written in natural languages (English, Spanish) and artificial languages (computer software) based on a simple expression for the entropy as a function of message length and specific word diversity. Code text…

计算与语言 · 计算机科学 2015-12-03 Gerardo Febres , Klaus Jaffe , Carlos Gershenson
‹ 上一页 1 2 3 10 下一页 ›