蒸馏数据复杂度如何影响非自回归机器翻译的质量与置信度
计算与语言
2021-05-28 v1
摘要
尽管非自回归(NAR)模型在机器翻译中展现出巨大前景,其应用受限于对自回归模型知识蒸馏的依赖。为解决此问题,我们试图理解蒸馏为何如此有效。先前工作表明蒸馏训练数据比人工翻译复杂度更低。基于 Levenshtein Transformer 与 Mask-Predict NAR 模型在 WMT14 德英任务上的实验,本文表明不同类型复杂度具有不同影响:降低词汇多样性与减小重排序复杂度均有助于 NAR 学习更好的源-目标对齐,从而提升翻译质量,而词汇多样性是蒸馏提升模型置信度的主要原因,其以不同方式影响不同 NAR 模型的校准。
引用
@article{arxiv.2105.12900,
title = {How Does Distilled Data Complexity Impact the Quality and Confidence of Non-Autoregressive Machine Translation?},
author = {Weijia Xu and Shuming Ma and Dongdong Zhang and Marine Carpuat},
journal= {arXiv preprint arXiv:2105.12900},
year = {2021}
}
备注
Findings of ACL 2021