利用多语言命名规范与语言无关文本特征作为跨语言文本分析的通用语言
计算与语言
2007-05-23 v1 信息检索
摘要
我们提出一种简单却高效的基本方法,用于多个多语言和跨语言语言技术应用,这些应用不局限于通常的两种或三种语言,而是可以 relatively 较少 effort 的方式扩展到更多语言。该方法包括:利用现有的多语言语言资源(如同义词词典、命名规范和地理命名数据库),以及利用存在的更或 less 语言无关文本项目(如日期、货币表达式、数字、姓名和 cognates)。将文本映射到多语言资源并识别不同语言文本之间词标记的链接是应用(如跨语言文档相似度计算、多语言聚类和分类、跨语言文档检索以及提供跨语言信息访问工具)的基本要素。
引用
@article{arxiv.cs/0609064,
title = {Exploiting multilingual nomenclatures and language-independent text features as an interlingua for cross-lingual text analysis applications},
author = {Ralf Steinberger and Bruno Pouliquen and Camelia Ignat},
journal= {arXiv preprint arXiv:cs/0609064},
year = {2007}
}
备注
The approach described in this paper is used to link related documents across languages in the multilingual news analysis system NewsExplorer, which is freely accessible at http://press.jrc.it/NewsExplorer . 11 pages