Word Sense Disambiguation using Optimised Combinations of Knowledge Sources
Abstract
Word sense disambiguation algorithms, with few exceptions, have made use of only one lexical knowledge source. We describe a system which performs unrestricted word sense disambiguation (on all content words in free text) by combining different knowledge sources: semantic preferences, dictionary definitions and subject/domain codes along with part-of-speech tags. The usefulness of these sources is optimised by means of a learning algorithm. We also describe the creation of a new sense tagged corpus by combining existing resources. Tested accuracy of our approach on this corpus exceeds 92%, demonstrating the viability of all-word disambiguation rather than restricting oneself to a small sample.
Cite
@article{arxiv.cmp-lg/9806014,
title = {Word Sense Disambiguation using Optimised Combinations of Knowledge Sources},
author = {Yorick Wilks and Mark Stevenson},
journal= {arXiv preprint arXiv:cmp-lg/9806014},
year = {2007}
}
Comments
7 pages, uses colacl.sty. To appear in the Proceedings of COLING-ACL '98