English

Large-Scale Machine Translation between Arabic and Hebrew: Available Corpora and Initial Results

Computation and Language 2016-09-27 v1

Abstract

Machine translation between Arabic and Hebrew has so far been limited by a lack of parallel corpora, despite the political and cultural importance of this language pair. Previous work relied on manually-crafted grammars or pivoting via English, both of which are unsatisfactory for building a scalable and accurate MT system. In this work, we compare standard phrase-based and neural systems on Arabic-Hebrew translation. We experiment with tokenization by external tools and sub-word modeling by character-level neural models, and show that both methods lead to improved translation performance, with a small advantage to the neural models.

Keywords

Cite

@article{arxiv.1609.07701,
  title  = {Large-Scale Machine Translation between Arabic and Hebrew: Available Corpora and Initial Results},
  author = {Yonatan Belinkov and James Glass},
  journal= {arXiv preprint arXiv:1609.07701},
  year   = {2016}
}

Comments

SeMaT 2016

R2 v1 2026-06-22T16:00:20.660Z