中文

重新审视数据泄漏对依存句法分析的影响

计算与语言 2022-03-25 v1

摘要

Søgaard (2020)的近期研究表明,除树库规模外,训练图与测试图之间的重叠(称为泄漏)比其他解释更能说明依存句法分析性能观测差异。本文重新审视该主张,在更多模型和语言上测试。我们发现其仅适用于零样本跨语言设置。随后我们提出一种更细粒度的此类泄漏度量,与原始度量不同,它不仅能解释而且与观测性能变化相关。代码和数据见:https://github.com/miriamwanner/reu-nlp-project

关键词

引用

@article{arxiv.2203.12815,
  title  = {Revisiting the Effects of Leakage on Dependency Parsing},
  author = {Nathaniel Krasner and Miriam Wanner and Antonios Anastasopoulos},
  journal= {arXiv preprint arXiv:2203.12815},
  year   = {2022}
}

备注

to be presented at ACL'22 Findings