English

Similarity-Based Methods For Word Sense Disambiguation

cmp-lg 2008-02-03 v1 Computation and Language

Abstract

We compare four similarity-based estimation methods against back-off and maximum-likelihood estimation methods on a pseudo-word sense disambiguation task in which we controlled for both unigram and bigram frequency. The similarity-based methods perform up to 40% better on this particular task. We also conclude that events that occur only once in the training set have major impact on similarity-based estimates.

Keywords

Cite

@article{arxiv.cmp-lg/9708010,
  title  = {Similarity-Based Methods For Word Sense Disambiguation},
  author = {Ido Dagan and Lillian Lee and Fernando Pereira},
  journal= {arXiv preprint arXiv:cmp-lg/9708010},
  year   = {2008}
}

Comments

7 pages, uses psfig.tex and aclap.sty

R2 v1 2026-07-22T09:58:48.233Z