MLSUM:多语言摘要语料库
计算与语言
2020-05-01 v1
摘要
我们提出 MLSUM,首个大规模多语言摘要(MultiLingual SUMmarization)数据集。它取自在线报纸,包含五种不同语言(即法语、德语、西班牙语、俄语、土耳其语)的 150 万篇以上文章/摘要对。结合来自流行的 CNN/Daily Mail 数据集的英文报纸,所收集的数据构成了一个大规模多语言数据集,可启文本摘要领域的新研究方向。我们基于最先进(SOTA)系统给出了跨语言对比分析。这些分析凸显了现有偏差,从而论证了使用多语言数据集的必要性。
引用
@article{arxiv.2004.14900,
title = {MLSUM: The Multilingual Summarization Corpus},
author = {Thomas Scialom and Paul-Alexis Dray and Sylvain Lamprier and Benjamin Piwowarski and Jacopo Staiano},
journal= {arXiv preprint arXiv:2004.14900},
year = {2020}
}