In this paper, we offer a preliminary investigation into the task of in-image machine translation: transforming an image containing text in one language into an image containing the same text in another language. We propose an end-to-end neural model for this task inspired by recent approaches to neural machine translation, and demonstrate promising initial results based purely on pixel-level supervision. We then offer a quantitative and qualitative evaluation of our system outputs and discuss some common failure modes. Finally, we conclude with directions for future work.
@article{arxiv.2010.10648,
title = {Towards End-to-End In-Image Neural Machine Translation},
author = {Elman Mansimov and Mitchell Stern and Mia Chen and Orhan Firat and Jakob Uszkoreit and Puneet Jain},
journal= {arXiv preprint arXiv:2010.10648},
year = {2020}
}
Comments
Accepted as an oral presentation at EMNLP, NLP Beyond Text workshop, 2020