English

MTet: Multi-domain Translation for English and Vietnamese

Computation and Language 2022-10-20 v2 Artificial Intelligence

Abstract

We introduce MTet, the largest publicly available parallel corpus for English-Vietnamese translation. MTet consists of 4.2M high-quality training sentence pairs and a multi-domain test set refined by the Vietnamese research community. Combining with previous works on English-Vietnamese translation, we grow the existing parallel dataset to 6.2M sentence pairs. We also release the first pretrained model EnViT5 for English and Vietnamese languages. Combining both resources, our model significantly outperforms previous state-of-the-art results by up to 2 points in translation BLEU score, while being 1.6 times smaller.

Keywords

Cite

@article{arxiv.2210.05610,
  title  = {MTet: Multi-domain Translation for English and Vietnamese},
  author = {Chinh Ngo and Trieu H. Trinh and Long Phan and Hieu Tran and Tai Dang and Hieu Nguyen and Minh Nguyen and Minh-Thang Luong},
  journal= {arXiv preprint arXiv:2210.05610},
  year   = {2022}
}