用于序列到序列学习的经典结构化预测损失
计算与语言
2018-10-09 v5
摘要
近期许多工作致力于在序列级别训练神经注意力模型,所用方法或为强化学习式方法,或为优化束搜索。在本文中,我们考察一系列广泛用于训练线性模型进行结构化预测的古典目标函数,并将其应用于神经序列到序列模型。我们的实验表明,在同等设置下,这些损失函数表现得出奇地好,略优于束搜索优化。我们还报告了在 IWSLT'14 德英翻译以及 Gigaword 抽象摘要任务上的新 state of the art 结果。在更大的 WMT'14 英法翻译任务上,序列级训练取得了 41.5 BLEU,与 state of the art 持平。
引用
@article{arxiv.1711.04956,
title = {Classical Structured Prediction Losses for Sequence to Sequence Learning},
author = {Sergey Edunov and Myle Ott and Michael Auli and David Grangier and Marc'Aurelio Ranzato},
journal= {arXiv preprint arXiv:1711.04956},
year = {2018}
}
备注
10 pages, NAACL 2018