English

Natural Language Processing in Ethiopian Languages: Current State, Challenges, and Opportunities

Computation and Language 2023-03-28 v1

Abstract

This survey delves into the current state of natural language processing (NLP) for four Ethiopian languages: Amharic, Afaan Oromo, Tigrinya, and Wolaytta. Through this paper, we identify key challenges and opportunities for NLP research in Ethiopia. Furthermore, we provide a centralized repository on GitHub that contains publicly available resources for various NLP tasks in these languages. This repository can be updated periodically with contributions from other researchers. Our objective is to identify research gaps and disseminate the information to NLP researchers interested in Ethiopian languages and encourage future research in this domain.

Cite

@article{arxiv.2303.14406,
  title  = {Natural Language Processing in Ethiopian Languages: Current State, Challenges, and Opportunities},
  author = {Atnafu Lambebo Tonja and Tadesse Destaw Belay and Israel Abebe Azime and Abinew Ali Ayele and Moges Ahmed Mehamed and Olga Kolesnikova and Seid Muhie Yimam},
  journal= {arXiv preprint arXiv:2303.14406},
  year   = {2023}
}

Comments

Accepted to Fourth workshop on Resources for African Indigenous Languages (RAIL), EACL2023

R2 v1 2026-06-28T09:33:20.432Z