English

RuBQ: A Russian Dataset for Question Answering over Wikidata

Computation and Language 2021-10-14 v1

Abstract

The paper presents RuBQ, the first Russian knowledge base question answering (KBQA) dataset. The high-quality dataset consists of 1,500 Russian questions of varying complexity, their English machine translations, SPARQL queries to Wikidata, reference answers, as well as a Wikidata sample of triples containing entities with Russian labels. The dataset creation started with a large collection of question-answer pairs from online quizzes. The data underwent automatic filtering, crowd-assisted entity linking, automatic generation of SPARQL queries, and their subsequent in-house verification.

Cite

@article{arxiv.2005.10659,
  title  = {RuBQ: A Russian Dataset for Question Answering over Wikidata},
  author = {Vladislav Korablinov and Pavel Braslavski},
  journal= {arXiv preprint arXiv:2005.10659},
  year   = {2021}
}
R2 v1 2026-06-23T15:43:01.126Z