English

Utilizing Lexical Similarity between Related, Low-resource Languages for Pivot-based SMT

Computation and Language 2017-10-06 v2

Abstract

We investigate pivot-based translation between related languages in a low resource, phrase-based SMT setting. We show that a subword-level pivot-based SMT model using a related pivot language is substantially better than word and morpheme-level pivot models. It is also highly competitive with the best direct translation model, which is encouraging as no direct source-target training corpus is used. We also show that combining multiple related language pivot models can rival a direct translation model. Thus, the use of subwords as translation units coupled with multiple related pivot languages can compensate for the lack of a direct parallel corpus.

Keywords

Cite

@article{arxiv.1702.07203,
  title  = {Utilizing Lexical Similarity between Related, Low-resource Languages for Pivot-based SMT},
  author = {Anoop Kunchukuttan and Maulik Shah and Pradyot Prakash and Pushpak Bhattacharyya},
  journal= {arXiv preprint arXiv:1702.07203},
  year   = {2017}
}

Comments

Accepted at IJCNLP 2017, 7 pages, 7 tables

R2 v1 2026-06-22T18:26:24.579Z