句子语义相关的成因:一个文本相关性数据集与实证研究
计算与语言
2023-03-21 v4
摘要
两个语言单元的语义相关程度长期以来被认为是理解意义的基础。此外,自动判定相关性在问答和摘要等许多应用中都有用处。然而,由于缺少相关性数据集,先前的 NLP 工作主要关注相关性子集——语义相似性。在本文中,我们引入一个语义文本相关性数据集 STR-2022,包含 5500 对英文句子对,采用比较式标注框架手动标注,得到细粒度分数。我们表明人类对句子对相关性的直觉高度可靠,重复标注相关性为 0.84。我们利用该数据集探究句子语义相关的成因。我们还展示了 STR-2022 在评估句子表示自动方法及各类下游 NLP 任务中的效用。我们的数据集、数据说明和标注问卷可在 https://doi.org/10.5281/zenodo.7599667 获取。
引用
@article{arxiv.2110.04845,
title = {What Makes Sentences Semantically Related: A Textual Relatedness Dataset and Empirical Study},
author = {Mohamed Abdalla and Krishnapriya Vishnubhotla and Saif M. Mohammad},
journal= {arXiv preprint arXiv:2110.04845},
year = {2023}
}
备注
Accepted to EACL 2023; Our dataset, data statement, and annotation questionnaire can be found at: https://doi.org/10.5281/zenodo.7599667