NICT 用于 WMT18 平行语料过滤任务的语料过滤系统
计算与语言
2018-10-15 v2
摘要
本文介绍了 NICT 参与 WMT18 共享平行语料过滤任务的情况。组织者提供了作为 Paracrawl 项目一部分、从网络爬取的 10 亿词德语-英语语料。该语料过于嘈杂,无法构建可用的神经机器翻译(NMT)系统。利用 WMT18 共享新闻翻译任务的干净数据,我们设计了若干特征并训练了一个分类器,为噪声数据中的每对句子打分。最终,我们抽样了 1 亿词和 1000 万词并构建了相应的 NMT 系统。实证结果表明,我们在抽样数据上训练的 NMT 系统取得了有前景的性能。
引用
@article{arxiv.1809.07043,
title = {NICT's Corpus Filtering Systems for the WMT18 Parallel Corpus Filtering Task},
author = {Rui Wang and Benjamin Marie and Masao Utiyama and Eiichiro Sumita},
journal= {arXiv preprint arXiv:1809.07043},
year = {2018}
}
备注
Due to the policy of our institute, with the agreement of all of the author, we decide to withdraw this paper