中文

摘要数据集的现状与命运:综述

计算与语言 2025-02-12 v2

摘要

自动摘要因其 versatility 和在 various downstream 任务中的广泛应用而持续受到关注。尽管如此,我们发现 annotation 努力 largely be disjointed,且 lacks common terminology。这使得发现 existing resources 或 identify coherent research directions 具有挑战性。为此,我们 survey 大量工作,spanning 133 个数据集 in over 100 种语言,构建了一个 novel ontology 涵盖 sample properties、collection methods 和 distribution。With this ontology 我们 make key observations,包括 low-resource 语言缺乏可访问 high-quality 数据集,以及该领域对 news domain 和自动收集的 distant supervision 的过度依赖。最后,我们提供了一个 web interface,allow users interact and explore our ontology 和 dataset collection,以及用于 streamline future research 的 summarization data card 模板。

关键词

引用

@article{arxiv.2411.04585,
  title  = {The State and Fate of Summarization Datasets: A Survey},
  author = {Noam Dahan and Gabriel Stanovsky},
  journal= {arXiv preprint arXiv:2411.04585},
  year   = {2025}
}

备注

Accepted to NAACL 2025