标注器未能习得之处,恰为解析器之所急需
计算与语言
2021-04-05 v1 人工智能
摘要
我们提出对神经 UPOS 标注器的错误分析,以评估为何使用金标准标签对解析性能有巨大正向贡献,而使用预测 UPOS 标签要么损害性能,要么仅带来可忽略的改进。我们评估神经依存解析器隐式习得了哪些关于词类型的知识,以及这与标注器所犯错误如何关联,从而解释使用预测标签对解析器影响甚微。我们还简要分析了哪些上下文导致标注性能下降。随后我们基于标注器所犯错误对 UPOS 标签进行掩码,以剥离标注器正确与错误分类的 UPOS 标签各自的贡献以及标注错误的影响。
引用
@article{arxiv.2104.01083,
title = {What Taggers Fail to Learn, Parsers Need the Most},
author = {Mark Anderson and Carlos Gómez-Rodríguez},
journal= {arXiv preprint arXiv:2104.01083},
year = {2021}
}
备注
Due to be published in the proceedings of the 23rd Nordic Conference on Computational Linguistics (NoDaLiDa 2021). Previously rejected at the 2021 Conference of the European Chapter of the Association for Computational Linguistics (EACL 2021)