English

Aligning a Parallel English-Chinese Corpus Statistically with Lexical Criteria

cmp-lg 2016-08-31 v1 Computation and Language

Abstract

We describe our experience with automatic alignment of sentences in parallel English-Chinese texts. Our report concerns three related topics: (1) progress on the HKUST English-Chinese Parallel Bilingual Corpus; (2) experiments addressing the applicability of Gale & Church's length-based statistical method to the task of alignment involving a non-Indo-European language; and (3) an improved statistical method that also incorporates domain-specific lexical cues.

Keywords

Cite

@article{arxiv.cmp-lg/9406007,
  title  = {Aligning a Parallel English-Chinese Corpus Statistically with Lexical Criteria},
  author = {Dekai Wu},
  journal= {arXiv preprint arXiv:cmp-lg/9406007},
  year   = {2016}
}

Comments

8 pages, uuencoded compressed PostScript