English

Talking to myself: self-dialogues as data for conversational agents

Computation and Language 2018-09-20 v2 Artificial Intelligence

Abstract

Conversational agents are gaining popularity with the increasing ubiquity of smart devices. However, training agents in a data driven manner is challenging due to a lack of suitable corpora. This paper presents a novel method for gathering topical, unstructured conversational data in an efficient way: self-dialogues through crowd-sourcing. Alongside this paper, we include a corpus of 3.6 million words across 23 topics. We argue the utility of the corpus by comparing self-dialogues with standard two-party conversations as well as data from other corpora.

Keywords

Cite

@article{arxiv.1809.06641,
  title  = {Talking to myself: self-dialogues as data for conversational agents},
  author = {Joachim Fainberg and Ben Krause and Mihai Dobre and Marco Damonte and Emmanuel Kahembwe and Daniel Duma and Bonnie Webber and Federico Fancellu},
  journal= {arXiv preprint arXiv:1809.06641},
  year   = {2018}
}

Comments

5 pages, 5 pages appendix, 2 figures

R2 v1 2026-06-23T04:09:53.129Z