中文
相关论文

相关论文: Character Entropy in Modern and Historical Texts: …

200 篇论文

While the use of statistical physics methods to analyze large corpora has been useful to unveil many patterns in texts, no comprehensive investigation has been performed investigating the properties of statistical measurements across…

The statistical properties of letters frequencies in European literature texts are investigated. The determination of logarithmic dependence of letters sequence for one-language and two-language texts are examined. The pare of languages is…

The Voynich Manuscript (VMS) exhibits a script of uncertain origin whose grapheme sequences have resisted linguistic analysis. We present a systematic analysis of its grapheme sequences, revealing two complementary structural layers: a…

计算与语言 · 计算机科学 2026-04-23 Christophe Parisel

We study the entropy of Chinese and English texts, based on characters in case of Chinese texts and based on words for both languages. Significant differences are found between the languages and between different personal styles of debating…

计算与语言 · 计算机科学 2017-01-17 R. R. Xie , W. B. Deng , D. J. Wang , L. P. Csernai

Departing from the postulate that Voynich Manuscript is not a hoax but rather encodes authentic contents, our article presents an evolutionary algorithm which aims to find the most optimal mapping between voynichian glyphs and candidate…

计算与语言 · 计算机科学 2021-07-13 Daniel Devatman Hromada

Written language is complex. A written text can be considered an attempt to convey a meaningful message which ends up being constrained by language rules, context dependence and highly redundant in its use of resources. Despite all these…

计算与语言 · 计算机科学 2019-05-20 E. Estevez-Rams , A. Mesa Rodriguez , D. Estevez-Moya

There are different ways of measuring diversity in complex systems. In particular, in language, lexical diversity is characterized in terms of the type-token ratio and the word entropy. We here investigate both diversity metrics in six…

计算与语言 · 计算机科学 2025-07-16 Pablo Rosillo-Rodes , Maxi San Miguel , David Sanchez

Witnesses of medieval literary texts, preserved in manuscript, are layered objects , being almost exclusively copies of copies. This results in multiple and hard to distinguish linguistic strata -- the author's scripta interacting with the…

计算与语言 · 计算机科学 2018-02-06 Jean-Baptiste Camps

We evaluated the impact of changing the observation scale over the entropy measures for text descriptions. MIDI coded Music, computer code and two human natural languages were studied at the scale of characters, words, and at the…

信息论 · 计算机科学 2017-01-13 Gerardo Febres , Klaus Jaffe

Diacritics are orthographic marks that clarify pronunciation, distinguish similar words, or alter meaning. They play a central role in many writing systems, yet their impact on language technology has not been systematically quantified…

计算与语言 · 计算机科学 2026-03-31 Adi Cohen , Yuval Pinter

We present a unified quantitative analysis of the Currier A/B language distinction in the Voynich Manuscript, proceeding in two stages. First, we confirm that the distinction is genuine: a Beta-Binomial mixture model applied to…

密码学与安全 · 计算机科学 2026-05-06 Christophe Parisel

The Voynich Manuscript is a medieval book written in an unknown script. This paper studies the distribution of similarly spelled words in the Voynich Manuscript. It shows that the distribution of words within the manuscript is not…

计算与语言 · 计算机科学 2016-02-09 Torsten Timm

This paper experiments with frequency-based corpus similarity measures across 39 languages using a register prediction task. The goal is to quantify (i) the distance between different corpora from the same language and (ii) the homogeneity…

计算与语言 · 计算机科学 2022-06-10 Haipeng Li , Jonathan Dunn

We compared entropy for texts written in natural languages (English, Spanish) and artificial languages (computer software) based on a simple expression for the entropy as a function of message length and specific word diversity. Code text…

计算与语言 · 计算机科学 2015-12-03 Gerardo Febres , Klaus Jaffe , Carlos Gershenson

Stylometry is mostly applied to authorial style. Recently, researchers have begun investigating the style of characters, finding that the variation remains within authorial bounds. We address the stylistic distinctiveness of characters in…

计算与语言 · 计算机科学 2023-01-16 Artjoms Šeļa , Ben Nagy , Joanna Byszuk , Laura Hernández-Lorenzo , Botond Szemes , Maciej Eder

Tokenisation - "the process of splitting text into atomic parts" (Brezina & Timperley, 2017: 1) - is a crucial step for corpus linguistics, as it provides the basis for any applicable quantitative method (e.g. collocations) while ensuring…

计算与语言 · 计算机科学 2025-07-03 Matteo Di Cristofaro

Texts exhibit considerable stylistic variation. This paper reports an experiment where a corpus of documents (N= 75 000) is analyzed using various simple stylistic metrics. A subset (n = 1000) of the corpus has been previously assessed to…

cmp-lg · 计算机科学 2008-02-03 Jussi Karlgren

We measured entropy and symbolic diversity for English and Spanish texts including literature Nobel laureates and other famous authors. Entropy, symbol diversity and symbol frequency profiles were compared for these four groups. We also…

计算与语言 · 计算机科学 2017-01-17 Gerardo Febres , Klaus Jaffe

The complexity of a system description is a function of the entropy of its symbolic description. Prior to computing the entropy of the system description, an observation scale has to be assumed. In natural language texts, typical scales are…

信息论 · 计算机科学 2015-03-31 Gerardo Febres , Klaus Jaffe

The translation of written language has been known since the 3rd century BC; however, its necessity has become increasingly common in the information age. Today, many translators exist, based on encoder-decoder deep architectures,…

计算与语言 · 计算机科学 2025-11-18 Ronit D. Gross , Yanir Harel , Ido Kanter
‹ 上一页 1 2 3 10 下一页 ›