The paper presents RuBQ, the first Russian knowledge base question answering (KBQA) dataset. The high-quality dataset consists of 1,500 Russian questions of varying complexity, their English machine translations, SPARQL queries to Wikidata, reference answers, as well as a Wikidata sample of triples containing entities with Russian labels. The dataset creation started with a large collection of question-answer pairs from online quizzes. The data underwent automatic filtering, crowd-assisted entity linking, automatic generation of SPARQL queries, and their subsequent in-house verification.
Cite
@article{arxiv.2005.10659,
title = {RuBQ: A Russian Dataset for Question Answering over Wikidata},
author = {Vladislav Korablinov and Pavel Braslavski},
journal= {arXiv preprint arXiv:2005.10659},
year = {2021}
}