中文

PhoMT:面向越南语-英语机器翻译的高质量大规模基准数据集

计算与语言 2021-10-26 v1

摘要

我们介绍了一个包含 3.02M 句对的越南语-英语高质量大规模平行数据集,比基准越南语-英语机器翻译语料库 IWSLT15 多出 2.9M 句对。我们在该数据集上对比了强神经基线与知名自动翻译引擎的实验,并在自动与人工评测中发现:最佳性能通过对预训练序列到序列去噪自编码器 mBART 进行微调取得。据我们所知,这是首个大规模越南语-英语机器翻译研究。我们希望公开的数据集与研究能作为未来越南语-英语机器翻译研究与应用的起点。

关键词

引用

@article{arxiv.2110.12199,
  title  = {PhoMT: A High-Quality and Large-Scale Benchmark Dataset for Vietnamese-English Machine Translation},
  author = {Long Doan and Linh The Nguyen and Nguyen Luong Tran and Thai Hoang and Dat Quoc Nguyen},
  journal= {arXiv preprint arXiv:2110.12199},
  year   = {2021}
}

备注

To appear in Proceedings of EMNLP 2021 (main conference). The first three authors contribute equally to this work