English

Domain-specific MT for Low-resource Languages: The case of Bambara-French

Computation and Language 2021-04-02 v1

Abstract

Translating to and from low-resource languages is a challenge for machine translation (MT) systems due to a lack of parallel data. In this paper we address the issue of domain-specific MT for Bambara, an under-resourced Mande language spoken in Mali. We present the first domain-specific parallel dataset for MT of Bambara into and from French. We discuss challenges in working with small quantities of domain-specific data for a low-resource language and we present the results of machine learning experiments on this data.

Keywords

Cite

@article{arxiv.2104.00041,
  title  = {Domain-specific MT for Low-resource Languages: The case of Bambara-French},
  author = {Allahsera Auguste Tapo and Michael Leventhal and Sarah Luger and Christopher M. Homan and Marcos Zampieri},
  journal= {arXiv preprint arXiv:2104.00041},
  year   = {2021}
}