English

AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages

Computation and Language 2020-05-04 v1

Abstract

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article category classification datasets for 9 languages to evaluate the embeddings. We show that the IndicNLP embeddings significantly outperform publicly available pre-trained embedding on multiple evaluation tasks. We hope that the availability of the corpus will accelerate Indic NLP research. The resources are available at https://github.com/ai4bharat-indicnlp/indicnlp_corpus.

Keywords

Cite

@article{arxiv.2005.00085,
  title  = {AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages},
  author = {Anoop Kunchukuttan and Divyanshu Kakwani and Satish Golla and Gokul N. C. and Avik Bhattacharyya and Mitesh M. Khapra and Pratyush Kumar},
  journal= {arXiv preprint arXiv:2005.00085},
  year   = {2020}
}

Comments

7 pages, 8 tables, https://github.com/ai4bharat-indicnlp/indicnlp_corpus