Ubuntu对话语料库:用于非结构化多轮对话系统研究的大规模数据集
计算与语言
2016-07-26 v3 人工智能
机器学习
神经与进化计算
摘要
本文介绍了 Ubuntu Dialogue Corpus,一个包含近100万轮多轮对话的数据集,总计超过700万条话语和1亿词。这为基于神经语言模型、能利用大量无标签数据构建对话管理器的研究提供了独特资源。该数据集兼具 Dialog State Tracking Challenge 数据集的对话多轮特性,以及 Twitter 等微博服务的非结构化交互性质。我们还描述了两种适用于分析该数据集的神经学习架构,并提供了在选取最佳下一回应任务上的基准性能。
引用
@article{arxiv.1506.08909,
title = {The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems},
author = {Ryan Lowe and Nissan Pow and Iulian Serban and Joelle Pineau},
journal= {arXiv preprint arXiv:1506.08909},
year = {2016}
}
备注
SIGDIAL 2015. 10 pages, 5 figures. Update includes link to new version of the dataset, with some added features and bug fixes. See: https://github.com/rkadlec/ubuntu-ranking-dataset-creator