English

Igbo-English Machine Translation: An Evaluation Benchmark

Computation and Language 2020-04-03 v1 Machine Learning

Abstract

Although researchers and practitioners are pushing the boundaries and enhancing the capacities of NLP tools and methods, works on African languages are lagging. A lot of focus on well resourced languages such as English, Japanese, German, French, Russian, Mandarin Chinese etc. Over 97% of the world's 7000 languages, including African languages, are low resourced for NLP i.e. they have little or no data, tools, and techniques for NLP research. For instance, only 5 out of 2965, 0.19% authors of full text papers in the ACL Anthology extracted from the 5 major conferences in 2018 ACL, NAACL, EMNLP, COLING and CoNLL, are affiliated to African institutions. In this work, we discuss our effort toward building a standard machine translation benchmark dataset for Igbo, one of the 3 major Nigerian languages. Igbo is spoken by more than 50 million people globally with over 50% of the speakers are in southeastern Nigeria. Igbo is low resourced although there have been some efforts toward developing IgboNLP such as part of speech tagging and diacritic restoration

Keywords

Cite

@article{arxiv.2004.00648,
  title  = {Igbo-English Machine Translation: An Evaluation Benchmark},
  author = {Ignatius Ezeani and Paul Rayson and Ikechukwu Onyenwe and Chinedu Uchechukwu and Mark Hepple},
  journal= {arXiv preprint arXiv:2004.00648},
  year   = {2020}
}

Comments

4 pages

R2 v1 2026-06-23T14:35:52.180Z