English

Noun-phrase co-occurrence statistics for semi-automatic semantic lexicon construction

Computation and Language 2007-05-23 v1

Abstract

Generating semantic lexicons semi-automatically could be a great time saver, relative to creating them by hand. In this paper, we present an algorithm for extracting potential entries for a category from an on-line corpus, based upon a small set of exemplars. Our algorithm finds more correct terms and fewer incorrect ones than previous work in this area. Additionally, the entries that are generated potentially provide broader coverage of the category than would occur to an individual coding them by hand. Our algorithm finds many terms not included within Wordnet (many more than previous algorithms), and could be viewed as an ``enhancer'' of existing broad-coverage resources.

Keywords

Cite

@article{arxiv.cs/0008026,
  title  = {Noun-phrase co-occurrence statistics for semi-automatic semantic lexicon construction},
  author = {Brian Roark and Eugene Charniak},
  journal= {arXiv preprint arXiv:cs/0008026},
  year   = {2007}
}

Comments

7 pages, 1 figure, 5 tables

R2 v1 2026-07-22T12:18:21.124Z