摘要数据集的现状与命运:综述
计算与语言
2025-02-12 v2
摘要
自动摘要因其 versatility 和在 various downstream 任务中的广泛应用而持续受到关注。尽管如此,我们发现 annotation 努力 largely be disjointed,且 lacks common terminology。这使得发现 existing resources 或 identify coherent research directions 具有挑战性。为此,我们 survey 大量工作,spanning 133 个数据集 in over 100 种语言,构建了一个 novel ontology 涵盖 sample properties、collection methods 和 distribution。With this ontology 我们 make key observations,包括 low-resource 语言缺乏可访问 high-quality 数据集,以及该领域对 news domain 和自动收集的 distant supervision 的过度依赖。最后,我们提供了一个 web interface,allow users interact and explore our ontology 和 dataset collection,以及用于 streamline future research 的 summarization data card 模板。
引用
@article{arxiv.2411.04585,
title = {The State and Fate of Summarization Datasets: A Survey},
author = {Noam Dahan and Gabriel Stanovsky},
journal= {arXiv preprint arXiv:2411.04585},
year = {2025}
}
备注
Accepted to NAACL 2025