中文

近缘印度语言的自动识别:资源与实验

计算与语言 2018-03-28 v1

摘要

在本文中,我们讨论了为印度的 5 种近缘印度-雅利安语支语言(阿瓦德语、博杰普尔语、布拉吉语、印地语和摩揭语)开发自动语言识别系统的尝试。我们从各种资源中为这些语言编译了长度不一的可比语料库。我们详细讨论了这些语料库的创建方法。使用这些语料库,我们开发了一个语言识别系统,目前达到了 96.48% 的 SOTA 准确率。我们还利用这些语料库研究了这 5 种语言在词汇层面的相似性,这是关于这些语言亲密程度的首次基于数据的研究。

关键词

引用

@article{arxiv.1803.09405,
  title  = {Automatic Identification of Closely-related Indian Languages: Resources and Experiments},
  author = {Ritesh Kumar and Bornini Lahiri and Deepak Alok and Atul Kr. Ojha and Mayank Jain and Abdul Basit and Yogesh Dawer},
  journal= {arXiv preprint arXiv:1803.09405},
  year   = {2018}
}

备注

Paper accepted at the 4th Workshop in Indian Languages Data and Resources (WILDRE - 4), 11th edition of the Language Resources and Evaluation Conference (LREC - 2018), 7-12 May 2018, Miyazaki (Japan)