Unsupervised Learning of Word-Category Guessing Rules
cmp-lg
2008-02-03 v1 Computation and Language
Abstract
Words unknown to the lexicon present a substantial problem to part-of-speech tagging. In this paper we present a technique for fully unsupervised statistical acquisition of rules which guess possible parts-of-speech for unknown words. Three complementary sets of word-guessing rules are induced from the lexicon and a raw corpus: prefix morphological rules, suffix morphological rules and ending-guessing rules. The learning was performed on the Brown Corpus data and rule-sets, with a highly competitive performance, were produced and compared with the state-of-the-art.
Cite
@article{arxiv.cmp-lg/9604022,
title = {Unsupervised Learning of Word-Category Guessing Rules},
author = {Andrei Mikheev},
journal= {arXiv preprint arXiv:cmp-lg/9604022},
year = {2008}
}
Comments
8 pages, LaTeX (aclap.sty for ACL-96); Proceedings of ACL-96 Santa Cruz, USA; also see cmp-lg/9604025