AI4Bharat-IndicNLP语料库:印度语言的单语语料与词嵌入
计算与语言
2020-05-04 v1
摘要
我们呈现IndicNLP语料库,一个大规模、通用领域的语料,包含来自两个语系的10种印度语言的27亿词。我们分享了在这些语料上训练的预训练词嵌入。我们为9种语言创建了新闻文章类别分类数据集以评估这些嵌入。我们表明IndicNLP嵌入在多个评估任务上显著优于公开可用的预训练嵌入。我们希望该语料的可获得性将加速印度NLP研究。相关资源位于 https://github.com/ai4bharat-indicnlp/indicnlp_corpus。
引用
@article{arxiv.2005.00085,
title = {AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages},
author = {Anoop Kunchukuttan and Divyanshu Kakwani and Satish Golla and Gokul N. C. and Avik Bhattacharyya and Mitesh M. Khapra and Pratyush Kumar},
journal= {arXiv preprint arXiv:2005.00085},
year = {2020}
}
备注
7 pages, 8 tables, https://github.com/ai4bharat-indicnlp/indicnlp_corpus