中文

AI4Bharat-IndicNLP语料库:印度语言的单语语料与词嵌入

计算与语言 2020-05-04 v1

摘要

我们呈现IndicNLP语料库,一个大规模、通用领域的语料,包含来自两个语系的10种印度语言的27亿词。我们分享了在这些语料上训练的预训练词嵌入。我们为9种语言创建了新闻文章类别分类数据集以评估这些嵌入。我们表明IndicNLP嵌入在多个评估任务上显著优于公开可用的预训练嵌入。我们希望该语料的可获得性将加速印度NLP研究。相关资源位于 https://github.com/ai4bharat-indicnlp/indicnlp_corpus。

关键词

引用

@article{arxiv.2005.00085,
  title  = {AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages},
  author = {Anoop Kunchukuttan and Divyanshu Kakwani and Satish Golla and Gokul N. C. and Avik Bhattacharyya and Mitesh M. Khapra and Pratyush Kumar},
  journal= {arXiv preprint arXiv:2005.00085},
  year   = {2020}
}

备注

7 pages, 8 tables, https://github.com/ai4bharat-indicnlp/indicnlp_corpus