English

emrQA: A Large Corpus for Question Answering on Electronic Medical Records

Computation and Language 2018-09-05 v1

Abstract

We propose a novel methodology to generate domain-specific large-scale question answering (QA) datasets by re-purposing existing annotations for other NLP tasks. We demonstrate an instance of this methodology in generating a large-scale QA dataset for electronic medical records by leveraging existing expert annotations on clinical notes for various NLP tasks from the community shared i2b2 datasets. The resulting corpus (emrQA) has 1 million question-logical form and 400,000+ question-answer evidence pairs. We characterize the dataset and explore its learning potential by training baseline models for question to logical form and question to answer mapping.

Keywords

Cite

@article{arxiv.1809.00732,
  title  = {emrQA: A Large Corpus for Question Answering on Electronic Medical Records},
  author = {Anusri Pampari and Preethi Raghavan and Jennifer Liang and Jian Peng},
  journal= {arXiv preprint arXiv:1809.00732},
  year   = {2018}
}

Comments

Accepted at Conference on Empirical Methods in Natural Language Processing (EMNLP) 2018