用于构建数据驱动对话系统的可用语料库综述
计算与语言
2017-03-22 v3 人工智能
人机交互
机器学习
机器学习
摘要
在过去的十年中,语音和语言理解的若干领域因使用数据驱动模型而取得了实质性突破。在对话系统领域,这一趋势不太明显,且大多数实用系统仍通过大量的工程化和专家知识构建。然而,近期的若干结果表明数据驱动方法是可行且颇有前景的。为促进该领域的研究,我们对适用于对话系统数据驱动学习的公开可用数据集进行了广泛调研。我们讨论了这些数据集的重要特征、如何用于学习多样的对话策略以及它们的其他潜在用途。我们还考察了数据集间的迁移学习方法以及外部知识的使用。最后,我们讨论了针对学习目标的评估指标的恰当选择。
引用
@article{arxiv.1512.05742,
title = {A Survey of Available Corpora for Building Data-Driven Dialogue Systems},
author = {Iulian Vlad Serban and Ryan Lowe and Peter Henderson and Laurent Charlin and Joelle Pineau},
journal= {arXiv preprint arXiv:1512.05742},
year = {2017}
}
备注
56 pages including references and appendix, 5 tables and 1 figure; Under review for the Dialogue & Discourse journal. Update: paper has been rewritten and now includes several new datasets