K-vec: A New Approach for Aligning Parallel Texts
cmp-lg
2008-02-03 v1 Computation and Language
Abstract
Various methods have been proposed for aligning texts in two or more languages such as the Canadian Parliamentary Debates(Hansards). Some of these methods generate a bilingual lexicon as a by-product. We present an alternative alignment strategy which we call K-vec, that starts by estimating the lexicon. For example, it discovers that the English word "fisheries" is similar to the French "pe^ches" by noting that the distribution of "fisheries" in the English text is similar to the distribution of "pe^ches" in the French. K-vec does not depend on sentence boundaries.
Cite
@article{arxiv.cmp-lg/9407021,
title = {K-vec: A New Approach for Aligning Parallel Texts},
author = {Pascale Fung and Kenneth Church},
journal= {arXiv preprint arXiv:cmp-lg/9407021},
year = {2008}
}
Comments
7 pages, uuencoded, compressed PostScript; Proc. COLING-94