English

How Much Reading Does Reading Comprehension Require? A Critical Investigation of Popular Benchmarks

Computation and Language 2018-08-22 v2 Artificial Intelligence Machine Learning Machine Learning

Abstract

Many recent papers address reading comprehension, where examples consist of (question, passage, answer) tuples. Presumably, a model must combine information from both questions and passages to predict corresponding answers. However, despite intense interest in the topic, with hundreds of published papers vying for leaderboard dominance, basic questions about the difficulty of many popular benchmarks remain unanswered. In this paper, we establish sensible baselines for the bAbI, SQuAD, CBT, CNN, and Who-did-What datasets, finding that question- and passage-only models often perform surprisingly well. On 1414 out of 2020 bAbI tasks, passage-only models achieve greater than 50%50\% accuracy, sometimes matching the full model. Interestingly, while CBT provides 2020-sentence stories only the last is needed for comparably accurate prediction. By comparison, SQuAD and CNN appear better-constructed.

Keywords

Cite

@article{arxiv.1808.04926,
  title  = {How Much Reading Does Reading Comprehension Require? A Critical Investigation of Popular Benchmarks},
  author = {Divyansh Kaushik and Zachary C. Lipton},
  journal= {arXiv preprint arXiv:1808.04926},
  year   = {2018}
}

Comments

To appear in EMNLP 2018