English

Revisiting Tri-training of Dependency Parsers

Computation and Language 2023-10-18 v1

Abstract

We compare two orthogonal semi-supervised learning techniques, namely tri-training and pretrained word embeddings, in the task of dependency parsing. We explore language-specific FastText and ELMo embeddings and multilingual BERT embeddings. We focus on a low resource scenario as semi-supervised learning can be expected to have the most impact here. Based on treebank size and available ELMo models, we select Hungarian, Uyghur (a zero-shot language for mBERT) and Vietnamese. Furthermore, we include English in a simulated low-resource setting. We find that pretrained word embeddings make more effective use of unlabelled data than tri-training but that the two approaches can be successfully combined.

Keywords

Cite

@article{arxiv.2109.08122,
  title  = {Revisiting Tri-training of Dependency Parsers},
  author = {Joachim Wagner and Jennifer Foster},
  journal= {arXiv preprint arXiv:2109.08122},
  year   = {2023}
}

Comments

17 pages, 1 figure, to be published at EMNLP 2021