English

Models of Co-occurrence

cmp-lg 2007-05-23 v1 Computation and Language

Abstract

A model of co-occurrence in bitext is a boolean predicate that indicates whether a given pair of word tokens co-occur in corresponding regions of the bitext space. Co-occurrence is a precondition for the possibility that two tokens might be mutual translations. Models of co-occurrence are the glue that binds methods for mapping bitext correspondence with methods for estimating translation models into an integrated system for exploiting parallel texts. Different models of co-occurrence are possible, depending on the kind of bitext map that is available, the language-specific information that is available, and the assumptions made about the nature of translational equivalence. Although most statistical translation models are based on models of co-occurrence, modeling co-occurrence correctly is more difficult than it may at first appear.

Keywords

Cite

@article{arxiv.cmp-lg/9805003,
  title  = {Models of Co-occurrence},
  author = {I. Dan Melamed},
  journal= {arXiv preprint arXiv:cmp-lg/9805003},
  year   = {2007}
}
R2 v1 2026-07-22T09:58:53.977Z