English

K-vec: A New Approach for Aligning Parallel Texts

cmp-lg 2008-02-03 v1 Computation and Language

Abstract

Various methods have been proposed for aligning texts in two or more languages such as the Canadian Parliamentary Debates(Hansards). Some of these methods generate a bilingual lexicon as a by-product. We present an alternative alignment strategy which we call K-vec, that starts by estimating the lexicon. For example, it discovers that the English word "fisheries" is similar to the French "pe^ches" by noting that the distribution of "fisheries" in the English text is similar to the distribution of "pe^ches" in the French. K-vec does not depend on sentence boundaries.

Cite

@article{arxiv.cmp-lg/9407021,
  title  = {K-vec: A New Approach for Aligning Parallel Texts},
  author = {Pascale Fung and Kenneth Church},
  journal= {arXiv preprint arXiv:cmp-lg/9407021},
  year   = {2008}
}

Comments

7 pages, uuencoded, compressed PostScript; Proc. COLING-94