This paper describes the submissions of the "Marian" team to the WNMT 2018 shared task. We investigate combinations of teacher-student training, low-precision matrix products, auto-tuning and other methods to optimize the Transformer model on GPU and CPU. By further integrating these methods with the new averaging attention networks, a recently introduced faster Transformer variant, we create a number of high-quality, high-performance models on the GPU and CPU, dominating the Pareto frontier for this shared task.
@article{arxiv.1805.12096,
title = {Marian: Cost-effective High-Quality Neural Machine Translation in C++},
author = {Marcin Junczys-Dowmunt and Kenneth Heafield and Hieu Hoang and Roman Grundkiewicz and Anthony Aue},
journal= {arXiv preprint arXiv:1805.12096},
year = {2018}
}
Comments
System submission to the Workshop for Neural Machine Translation 2018, efficiency task