神经机器翻译与基于短语的机器翻译的细粒度人工评估
计算与语言
2018-02-13 v1
摘要
我们通过错误标注对系统输出进行细粒度人工评估,比较了三种统计机器翻译方法(纯基于短语、因子化基于短语和神经网络)。我们标注中的错误类型符合多维质量度量(MQM),且标注由两名标注员完成。在此类任务中标注员间的一致性很高,结果表明,表现最好的系统(神经网络)将表现最差的系统(基于短语)产生的错误减少了 54%。
引用
@article{arxiv.1706.04389,
title = {Fine-grained human evaluation of neural versus phrase-based machine translation},
author = {Filip Klubička and Antonio Toral and Víctor M. Sánchez-Cartagena},
journal= {arXiv preprint arXiv:1706.04389},
year = {2018}
}
备注
12 pages, 2 figures, The Prague Bulletin of Mathematical Linguistics