自然语言书面文本的词长熵与相关性
计算与语言
2014-01-27 v1 数据分析、统计与概率
摘要
我们研究了十种欧洲语言的词长频率分布及其相关性。研究结果表明:a) 由均值和熵量化的短词词长分布可将乌拉尔语系(芬兰语)语料库与其他语料库区分开来;b) 长词尾部(表现为分布的高阶矩)可将日耳曼语族(英语除外)与罗曼语族及希腊语区分开来;c) 通过比较真实熵与打乱文本的熵所测得的邻近词长之间的相关性,在日耳曼语族和芬兰语中较小。
引用
@article{arxiv.1401.6224,
title = {Word-length entropies and correlations of natural language written texts},
author = {Maria Kalimeri and Vassilios Constantoudis and Constantinos Papadimitriou and Konstantinos Karamanos and Fotis K. Diakonos and Harris Papageorgiou},
journal= {arXiv preprint arXiv:1401.6224},
year = {2014}
}
备注
13 pages + 1 page of supporting information, 9 figures