重新审视数据泄漏对依存句法分析的影响
计算与语言
2022-03-25 v1
摘要
Søgaard (2020)的近期研究表明,除树库规模外,训练图与测试图之间的重叠(称为泄漏)比其他解释更能说明依存句法分析性能观测差异。本文重新审视该主张,在更多模型和语言上测试。我们发现其仅适用于零样本跨语言设置。随后我们提出一种更细粒度的此类泄漏度量,与原始度量不同,它不仅能解释而且与观测性能变化相关。代码和数据见:https://github.com/miriamwanner/reu-nlp-project
引用
@article{arxiv.2203.12815,
title = {Revisiting the Effects of Leakage on Dependency Parsing},
author = {Nathaniel Krasner and Miriam Wanner and Antonios Anastasopoulos},
journal= {arXiv preprint arXiv:2203.12815},
year = {2022}
}
备注
to be presented at ACL'22 Findings