神经回复生成模型为何偏好通用回复?
计算与语言
2019-12-12 v2
摘要
序列到序列学习的最新进展揭示了面向回复生成任务的纯数据驱动方法。尽管应用广泛,现有神经模型易于产生简短且通用的回复,难以应对开放域挑战。本研究中,我们依据模型的优化目标及人-人对话语料的特定特征分析这一关键问题。通过将黑盒分解为若干部分,对概率极限进行详细分析以揭示这些通用回复背后的原因。基于上述分析,我们提出最大间隔排序正则项以避免模型偏向此类回复。最后,在案例研究与基准上以若干指标开展的实证实验验证了该方法。
引用
@article{arxiv.1808.09187,
title = {Why Do Neural Response Generation Models Prefer Universal Replies?},
author = {Bowen Wu and Nan Jiang and Zhifeng Gao and Mengyuan Li and Zongsheng Wang and Suke Li and Qihang Feng and Wenge Rong and Baoxun Wang},
journal= {arXiv preprint arXiv:1808.09187},
year = {2019}
}
备注
Preprint of the paper presented to the AAAI 2020 Workshop on Interactive and Conversational Recommendation Systems (WICRS)