English

Dataset for the First Evaluation on Chinese Machine Reading Comprehension

Computation and Language 2018-03-16 v2

Abstract

Machine Reading Comprehension (MRC) has become enormously popular recently and has attracted a lot of attention. However, existing reading comprehension datasets are mostly in English. To add diversity in reading comprehension datasets, in this paper we propose a new Chinese reading comprehension dataset for accelerating related research in the community. The proposed dataset contains two different types: cloze-style reading comprehension and user query reading comprehension, associated with large-scale training data as well as human-annotated validation and hidden test set. Along with this dataset, we also hosted the first Evaluation on Chinese Machine Reading Comprehension (CMRC-2017) and successfully attracted tens of participants, which suggest the potential impact of this dataset.

Keywords

Cite

@article{arxiv.1709.08299,
  title  = {Dataset for the First Evaluation on Chinese Machine Reading Comprehension},
  author = {Yiming Cui and Ting Liu and Zhipeng Chen and Wentao Ma and Shijin Wang and Guoping Hu},
  journal= {arXiv preprint arXiv:1709.08299},
  year   = {2018}
}

Comments

5 pages, published at LREC 2018

R2 v1 2026-06-22T21:53:19.858Z