人工序列与复杂度度量
统计力学
2009-11-10 v2 计算与语言
信息检索
信息论
math.IT
摘要
本文利用信息论概念来解决识别和定义最佳工具以自动且不偏向地从任意字符序列中提取信息的根本问题。我们特别引入一类方法,这些方法在关键方式上使用数据压缩技术,以基于其相对信息内容定义字符序列(例如文本)之间的距离和遥远程度度量。我们还详细讨论了如何利用数据压缩技术的特定特征来引入给定序列词典、人工文本的概念,并展示了这些新工具如何用于信息提取目的。我们指出了该方法的通用性和普遍性,这些方法适用于任何类型的字符词语集合,独立于其背后的编码类型。我们将语言动机问题作为案例研究,present for automatic language recognition, authorship attribution and self consistent-classification。
引用
@article{arxiv.cond-mat/0403233,
title = {Artificial Sequences and Complexity Measures},
author = {Andrea Baronchelli and Emanuele Caglioti and Vittorio Loreto},
journal= {arXiv preprint arXiv:cond-mat/0403233},
year = {2009}
}
备注
Revised version, with major changes, of previous "Data Compression approach to Information Extraction and Classification" by A. Baronchelli and V. Loreto. 15 pages; 5 figures