English

Error-Driven Pruning of Treebank Grammars for Base Noun Phrase Identification

cmp-lg 2007-05-23 v1 Computation and Language

Abstract

Finding simple, non-recursive, base noun phrases is an important subtask for many natural language processing applications. While previous empirical methods for base NP identification have been rather complex, this paper instead proposes a very simple algorithm that is tailored to the relative simplicity of the task. In particular, we present a corpus-based approach for finding base NPs by matching part-of-speech tag sequences. The training phase of the algorithm is based on two successful techniques: first the base NP grammar is read from a ``treebank'' corpus; then the grammar is improved by selecting rules with high ``benefit'' scores. Using this simple algorithm with a naive heuristic for matching rules, we achieve surprising accuracy in an evaluation on the Penn Treebank Wall Street Journal.

Keywords

Cite

@article{arxiv.cmp-lg/9808015,
  title  = {Error-Driven Pruning of Treebank Grammars for Base Noun Phrase Identification},
  author = {Claire Cardie and David Pierce},
  journal= {arXiv preprint arXiv:cmp-lg/9808015},
  year   = {2007}
}

Comments

7 pages; 2 eps figures; uses epsf, colacl

R2 v1 2026-07-22T09:58:59.443Z