英语机器阅读理解数据集综述
计算与语言
2021-10-11 v2
摘要
本文综述了60个英语机器阅读理解数据集,旨在为其他关注该问题的研究者提供便捷的资源。我们根据数据集的问题与答案形式对其进行分类,并从规模、词汇量、数据来源、创建方法、人类表现水平以及首个疑问词等多个维度进行比较。我们的分析表明,维基百科迄今为止是最常见的数据来源,并且各数据集中相对缺乏 why、when 和 where 类问题。
引用
@article{arxiv.2101.10421,
title = {English Machine Reading Comprehension Datasets: A Survey},
author = {Daria Dzendzik and Carl Vogel and Jennifer Foster},
journal= {arXiv preprint arXiv:2101.10421},
year = {2021}
}
备注
Will appear at EMNLP 2021. Dataset survey paper: 9 pages, 5 figures, 2 tables + attachment