English

Better Neural Machine Translation by Extracting Linguistic Information from BERT

Computation and Language 2021-04-08 v1

Abstract

Adding linguistic information (syntax or semantics) to neural machine translation (NMT) has mostly focused on using point estimates from pre-trained models. Directly using the capacity of massive pre-trained contextual word embedding models such as BERT (Devlin et al., 2019) has been marginally useful in NMT because effective fine-tuning is difficult to obtain for NMT without making training brittle and unreliable. We augment NMT by extracting dense fine-tuned vector-based linguistic information from BERT instead of using point estimates. Experimental results show that our method of incorporating linguistic information helps NMT to generalize better in a variety of training contexts and is no more difficult to train than conventional Transformer-based NMT.

Keywords

Cite

@article{arxiv.2104.02831,
  title  = {Better Neural Machine Translation by Extracting Linguistic Information from BERT},
  author = {Hassan S. Shavarani and Anoop Sarkar},
  journal= {arXiv preprint arXiv:2104.02831},
  year   = {2021}
}
R2 v1 2026-06-24T00:54:25.419Z