English

How Do Source-side Monolingual Word Embeddings Impact Neural Machine Translation?

Computation and Language 2018-06-15 v2

Abstract

Using pre-trained word embeddings as input layer is a common practice in many natural language processing (NLP) tasks, but it is largely neglected for neural machine translation (NMT). In this paper, we conducted a systematic analysis on the effect of using pre-trained source-side monolingual word embedding in NMT. We compared several strategies, such as fixing or updating the embeddings during NMT training on varying amounts of data, and we also proposed a novel strategy called dual-embedding that blends the fixing and updating strategies. Our results suggest that pre-trained embeddings can be helpful if properly incorporated into NMT, especially when parallel data is limited or additional in-domain monolingual data is readily available.

Keywords

Cite

@article{arxiv.1806.01515,
  title  = {How Do Source-side Monolingual Word Embeddings Impact Neural Machine Translation?},
  author = {Shuoyang Ding and Kevin Duh},
  journal= {arXiv preprint arXiv:1806.01515},
  year   = {2018}
}

Comments

10 pages, 4 figures

R2 v1 2026-06-23T02:19:14.737Z