中文

自然语言书面文本的词长熵与相关性

计算与语言 2014-01-27 v1 数据分析、统计与概率

摘要

我们研究了十种欧洲语言的词长频率分布及其相关性。研究结果表明:a) 由均值和熵量化的短词词长分布可将乌拉尔语系(芬兰语)语料库与其他语料库区分开来;b) 长词尾部(表现为分布的高阶矩)可将日耳曼语族(英语除外)与罗曼语族及希腊语区分开来;c) 通过比较真实熵与打乱文本的熵所测得的邻近词长之间的相关性,在日耳曼语族和芬兰语中较小。

关键词

引用

@article{arxiv.1401.6224,
  title  = {Word-length entropies and correlations of natural language written texts},
  author = {Maria Kalimeri and Vassilios Constantoudis and Constantinos Papadimitriou and Konstantinos Karamanos and Fotis K. Diakonos and Harris Papageorgiou},
  journal= {arXiv preprint arXiv:1401.6224},
  year   = {2014}
}

备注

13 pages + 1 page of supporting information, 9 figures