English

Cross-Lingual Semantic Role Labeling with High-Quality Translated Training Corpus

Computation and Language 2020-05-08 v2

Abstract

Many efforts of research are devoted to semantic role labeling (SRL) which is crucial for natural language understanding. Supervised approaches have achieved impressing performances when large-scale corpora are available for resource-rich languages such as English. While for the low-resource languages with no annotated SRL dataset, it is still challenging to obtain competitive performances. Cross-lingual SRL is one promising way to address the problem, which has achieved great advances with the help of model transferring and annotation projection. In this paper, we propose a novel alternative based on corpus translation, constructing high-quality training datasets for the target languages from the source gold-standard SRL annotations. Experimental results on Universal Proposition Bank show that the translation-based method is highly effective, and the automatic pseudo datasets can improve the target-language SRL performances significantly.

Keywords

Cite

@article{arxiv.2004.06295,
  title  = {Cross-Lingual Semantic Role Labeling with High-Quality Translated Training Corpus},
  author = {Hao Fei and Meishan Zhang and Donghong Ji},
  journal= {arXiv preprint arXiv:2004.06295},
  year   = {2020}
}

Comments

Accepted at ACL 2020

R2 v1 2026-06-23T14:50:15.671Z