English

Selecting Artificially-Generated Sentences for Fine-Tuning Neural Machine Translation

Computation and Language 2019-09-27 v1

Abstract

Neural Machine Translation (NMT) models tend to achieve best performance when larger sets of parallel sentences are provided for training. For this reason, augmenting the training set with artificially-generated sentence pairs can boost performance. Nonetheless, the performance can also be improved with a small number of sentences if they are in the same domain as the test set. Accordingly, we want to explore the use of artificially-generated sentences along with data-selection algorithms to improve German-to-English NMT models trained solely with authentic data. In this work, we show how artificially-generated sentences can be more beneficial than authentic pairs, and demonstrate their advantages when used in combination with data-selection algorithms.

Keywords

Cite

@article{arxiv.1909.12016,
  title  = {Selecting Artificially-Generated Sentences for Fine-Tuning Neural Machine Translation},
  author = {Alberto Poncelas and Andy Way},
  journal= {arXiv preprint arXiv:1909.12016},
  year   = {2019}
}

Comments

Proceedings of the 12th International Conference on Natural Language Generation (INLG 2019)

R2 v1 2026-06-23T11:26:43.725Z