English

Chinese-Japanese Unsupervised Neural Machine Translation Using Sub-character Level Information

Computation and Language 2019-03-04 v1

Abstract

Unsupervised neural machine translation (UNMT) requires only monolingual data of similar language pairs during training and can produce bi-directional translation models with relatively good performance on alphabetic languages (Lample et al., 2018). However, no research has been done to logographic language pairs. This study focuses on Chinese-Japanese UNMT trained by data containing sub-character (ideograph or stroke) level information which is decomposed from character level data. BLEU scores of both character and sub-character level systems were compared against each other and the results showed that despite the effectiveness of UNMT on character level data, sub-character level data could further enhance the performance, in which the stroke level system outperformed the ideograph level system.

Keywords

Cite

@article{arxiv.1903.00149,
  title  = {Chinese-Japanese Unsupervised Neural Machine Translation Using Sub-character Level Information},
  author = {Longtu Zhang and Mamoru Komachi},
  journal= {arXiv preprint arXiv:1903.00149},
  year   = {2019}
}

Comments

5 pages

R2 v1 2026-06-23T07:55:02.365Z