English

Integrating Multiple Knowledge Sources to Disambiguate Word Sense: An Exemplar-Based Approach

cmp-lg 2008-02-03 v1 Computation and Language

Abstract

In this paper, we present a new approach for word sense disambiguation (WSD) using an exemplar-based learning algorithm. This approach integrates a diverse set of knowledge sources to disambiguate word sense, including part of speech of neighboring words, morphological form, the unordered set of surrounding words, local collocations, and verb-object syntactic relation. We tested our WSD program, named {\sc Lexas}, on both a common data set used in previous work, as well as on a large sense-tagged corpus that we separately constructed. {\sc Lexas} achieves a higher accuracy on the common data set, and performs better than the most frequent heuristic on the highly ambiguous words in the large corpus tagged with the refined senses of {\sc WordNet}.

Keywords

Cite

@article{arxiv.cmp-lg/9606032,
  title  = {Integrating Multiple Knowledge Sources to Disambiguate Word Sense: An Exemplar-Based Approach},
  author = {Hwee Tou Ng and Hian Beng Lee},
  journal= {arXiv preprint arXiv:cmp-lg/9606032},
  year   = {2008}
}

Comments

In Proceedings of ACL96, 8 pages

R2 v1 2026-07-22T09:58:31.445Z