English

Plan, Attend, Generate: Character-level Neural Machine Translation with Planning in the Decoder

Computation and Language 2017-06-26 v2 Neural and Evolutionary Computing

Abstract

We investigate the integration of a planning mechanism into an encoder-decoder architecture with an explicit alignment for character-level machine translation. We develop a model that plans ahead when it computes alignments between the source and target sequences, constructing a matrix of proposed future alignments and a commitment vector that governs whether to follow or recompute the plan. This mechanism is inspired by the strategic attentive reader and writer (STRAW) model. Our proposed model is end-to-end trainable with fully differentiable operations. We show that it outperforms a strong baseline on three character-level decoder neural machine translation on WMT'15 corpus. Our analysis demonstrates that our model can compute qualitatively intuitive alignments and achieves superior performance with fewer parameters.

Keywords

Cite

@article{arxiv.1706.05087,
  title  = {Plan, Attend, Generate: Character-level Neural Machine Translation with Planning in the Decoder},
  author = {Caglar Gulcehre and Francis Dutil and Adam Trischler and Yoshua Bengio},
  journal= {arXiv preprint arXiv:1706.05087},
  year   = {2017}
}

Comments

Accepted to Rep4NLP 2017 Workshop at ACL 2017 Conference

R2 v1 2026-06-22T20:20:22.278Z