English

KGConv, a Conversational Corpus grounded in Wikidata

Computation and Language 2023-08-30 v1 Artificial Intelligence

Abstract

We present KGConv, a large, conversational corpus of 71k conversations where each question-answer pair is grounded in a Wikidata fact. Conversations contain on average 8.6 questions and for each Wikidata fact, we provide multiple variants (12 on average) of the corresponding question using templates, human annotations, hand-crafted rules and a question rewriting neural model. We provide baselines for the task of Knowledge-Based, Conversational Question Generation. KGConv can further be used for other generation and analysis tasks such as single-turn question generation from Wikidata triples, question rewriting, question answering from conversation or from knowledge graphs and quiz generation.

Cite

@article{arxiv.2308.15298,
  title  = {KGConv, a Conversational Corpus grounded in Wikidata},
  author = {Quentin Brabant and Gwenole Lecorve and Lina M. Rojas-Barahona and Claire Gardent},
  journal= {arXiv preprint arXiv:2308.15298},
  year   = {2023}
}
R2 v1 2026-06-28T12:07:21.706Z