English

Neural Machine Translation for Low-Resource Languages: A Survey

Computation and Language 2021-06-30 v1 Artificial Intelligence

Abstract

Neural Machine Translation (NMT) has seen a tremendous spurt of growth in less than ten years, and has already entered a mature phase. While considered as the most widely used solution for Machine Translation, its performance on low-resource language pairs still remains sub-optimal compared to the high-resource counterparts, due to the unavailability of large parallel corpora. Therefore, the implementation of NMT techniques for low-resource language pairs has been receiving the spotlight in the recent NMT research arena, thus leading to a substantial amount of research reported on this topic. This paper presents a detailed survey of research advancements in low-resource language NMT (LRL-NMT), along with a quantitative analysis aimed at identifying the most popular solutions. Based on our findings from reviewing previous work, this survey paper provides a set of guidelines to select the possible NMT technique for a given LRL data setting. It also presents a holistic view of the LRL-NMT research landscape and provides a list of recommendations to further enhance the research efforts on LRL-NMT.

Keywords

Cite

@article{arxiv.2106.15115,
  title  = {Neural Machine Translation for Low-Resource Languages: A Survey},
  author = {Surangika Ranathunga and En-Shiun Annie Lee and Marjana Prifti Skenduli and Ravi Shekhar and Mehreen Alam and Rishemjit Kaur},
  journal= {arXiv preprint arXiv:2106.15115},
  year   = {2021}
}

Comments

35 pages, 8 figures

R2 v1 2026-06-24T03:42:00.575Z