中文

面向英越语音翻译的高质量大规模数据集

计算与语言 2022-08-09 v1

摘要

本文介绍了一个面向英越语音翻译的高质量大规模基准数据集,包含 508 个音频小时,由 331K 个(句子长度音频、英文源转录句、越南语目标字幕句)三元组构成。我们还使用强基线进行了实证实验,发现传统的“级联(Cascaded)”方法仍优于现代的“端到端(End-to-End)”方法。据我们所知,这是首个大规模的英越语音翻译研究。我们希望我们公开的数据集与研究能作为未来英越语音翻译研究与应用的起点。我们的数据集可在 https://github.com/VinAIResearch/PhoST 获取。

关键词

引用

@article{arxiv.2208.04243,
  title  = {A High-Quality and Large-Scale Dataset for English-Vietnamese Speech Translation},
  author = {Linh The Nguyen and Nguyen Luong Tran and Long Doan and Manh Luong and Dat Quoc Nguyen},
  journal= {arXiv preprint arXiv:2208.04243},
  year   = {2022}
}

备注

In Proceedings of INTERSPEECH 2022, to appear. The first three authors contributed equally to this work