English

A Repository of Conversational Datasets

Computation and Language 2019-05-30 v2

Abstract

Progress in Machine Learning is often driven by the availability of large datasets, and consistent evaluation metrics for comparing modeling approaches. To this end, we present a repository of conversational datasets consisting of hundreds of millions of examples, and a standardised evaluation procedure for conversational response selection models using '1-of-100 accuracy'. The repository contains scripts that allow researchers to reproduce the standard datasets, or to adapt the pre-processing and data filtering steps to their needs. We introduce and evaluate several competitive baselines for conversational response selection, whose implementations are shared in the repository, as well as a neural encoder model that is trained on the entire training set.

Keywords

Cite

@article{arxiv.1904.06472,
  title  = {A Repository of Conversational Datasets},
  author = {Matthew Henderson and Paweł Budzianowski and Iñigo Casanueva and Sam Coope and Daniela Gerz and Girish Kumar and Nikola Mrkšić and Georgios Spithourakis and Pei-Hao Su and Ivan Vulić and Tsung-Hsien Wen},
  journal= {arXiv preprint arXiv:1904.06472},
  year   = {2019}
}
R2 v1 2026-06-23T08:38:31.369Z