English

Modeling Latent Sentence Structure in Neural Machine Translation

Computation and Language 2020-06-22 v2

Abstract

Recently it was shown that linguistic structure predicted by a supervised parser can be beneficial for neural machine translation (NMT). In this work we investigate a more challenging setup: we incorporate sentence structure as a latent variable in a standard NMT encoder-decoder and induce it in such a way as to benefit the translation task. We consider German-English and Japanese-English translation benchmarks and observe that when using RNN encoders the model makes no or very limited use of the structure induction apparatus. In contrast, CNN and word-embedding-based encoders rely on latent graphs and force them to encode useful, potentially long-distance, dependencies.

Keywords

Cite

@article{arxiv.1901.06436,
  title  = {Modeling Latent Sentence Structure in Neural Machine Translation},
  author = {Jasmijn Bastings and Wilker Aziz and Ivan Titov and Khalil Sima'an},
  journal= {arXiv preprint arXiv:1901.06436},
  year   = {2020}
}

Comments

Accepted as an extended abstract to ACL NMT workshop 2018

R2 v1 2026-06-23T07:16:15.523Z