English

KorQuAD1.0: Korean QA Dataset for Machine Reading Comprehension

Computation and Language 2019-09-18 v2

Abstract

Machine Reading Comprehension (MRC) is a task that requires machine to understand natural language and answer questions by reading a document. It is the core of automatic response technology such as chatbots and automatized customer supporting systems. We present Korean Question Answering Dataset(KorQuAD), a large-scale Korean dataset for extractive machine reading comprehension task. It consists of 70,000+ human generated question-answer pairs on Korean Wikipedia articles. We release KorQuAD1.0 and launch a challenge at https://KorQuAD.github.io to encourage the development of multilingual natural language processing research.

Keywords

Cite

@article{arxiv.1909.07005,
  title  = {KorQuAD1.0: Korean QA Dataset for Machine Reading Comprehension},
  author = {Seungyoung Lim and Myungji Kim and Jooyoul Lee},
  journal= {arXiv preprint arXiv:1909.07005},
  year   = {2019}
}
R2 v1 2026-06-23T11:16:13.232Z