English

An Exploration of Data Augmentation Techniques for Improving English to Tigrinya Translation

Computation and Language 2021-04-06 v2

Abstract

It has been shown that the performance of neural machine translation (NMT) drops starkly in low-resource conditions, often requiring large amounts of auxiliary data to achieve competitive results. An effective method of generating auxiliary data is back-translation of target language sentences. In this work, we present a case study of Tigrinya where we investigate several back-translation methods to generate synthetic source sentences. We find that in low-resource conditions, back-translation by pivoting through a higher-resource language related to the target language proves most effective resulting in substantial improvements over baselines.

Keywords

Cite

@article{arxiv.2103.16789,
  title  = {An Exploration of Data Augmentation Techniques for Improving English to Tigrinya Translation},
  author = {Lidia Kidane and Sachin Kumar and Yulia Tsvetkov},
  journal= {arXiv preprint arXiv:2103.16789},
  year   = {2021}
}

Comments

Accepted at AfricaNLP Workshop, EACL 2021