中文

Ubuntu对话语料库:用于非结构化多轮对话系统研究的大规模数据集

计算与语言 2016-07-26 v3 人工智能 机器学习 神经与进化计算

摘要

本文介绍了 Ubuntu Dialogue Corpus,一个包含近100万轮多轮对话的数据集,总计超过700万条话语和1亿词。这为基于神经语言模型、能利用大量无标签数据构建对话管理器的研究提供了独特资源。该数据集兼具 Dialog State Tracking Challenge 数据集的对话多轮特性,以及 Twitter 等微博服务的非结构化交互性质。我们还描述了两种适用于分析该数据集的神经学习架构,并提供了在选取最佳下一回应任务上的基准性能。

关键词

引用

@article{arxiv.1506.08909,
  title  = {The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems},
  author = {Ryan Lowe and Nissan Pow and Iulian Serban and Joelle Pineau},
  journal= {arXiv preprint arXiv:1506.08909},
  year   = {2016}
}

备注

SIGDIAL 2015. 10 pages, 5 figures. Update includes link to new version of the dataset, with some added features and bug fixes. See: https://github.com/rkadlec/ubuntu-ranking-dataset-creator