端到端关系抽取评估中保留与抽取的分离
计算与语言
2021-09-27 v1 机器学习
摘要
最先进的 NLP 模型可能采用浅层启发式方法,从而限制其泛化能力(McCoy et al., 2019)。此类启发式包括在命名实体识别中与训练集的词法重叠(Taillé et al., 2020)以及关系抽取中的事件或类型启发式(Rosenman et al., 2020)。在更贴近实际的端到端关系抽取(RE)设定中,我们可以预期另一种启发式:仅仅保留训练中的关系三元组。在本文中,我们提出了若干实验,证实已知事实的保留是标准基准上性能的一个关键因素。此外,一个实验表明,能够使用中间类型表示的流水线模型较不易过度依赖保留。
引用
@article{arxiv.2109.12008,
title = {Separating Retention from Extraction in the Evaluation of End-to-end Relation Extraction},
author = {Bruno Taillé and Vincent Guigue and Geoffrey Scoutheeten and Patrick Gallinari},
journal= {arXiv preprint arXiv:2109.12008},
year = {2021}
}
备注
Accepted at EMNLP 2021