English

WebSets: Extracting Sets of Entities from the Web Using Unsupervised Information Extraction

Machine Learning 2013-07-02 v1 Computation and Language Information Retrieval

Abstract

We describe a open-domain information extraction method for extracting concept-instance pairs from an HTML corpus. Most earlier approaches to this problem rely on combining clusters of distributionally similar terms and concept-instance pairs obtained with Hearst patterns. In contrast, our method relies on a novel approach for clustering terms found in HTML tables, and then assigning concept names to these clusters using Hearst patterns. The method can be efficiently applied to a large corpus, and experimental results on several datasets show that our method can accurately extract large numbers of concept-instance pairs.

Keywords

Cite

@article{arxiv.1307.0261,
  title  = {WebSets: Extracting Sets of Entities from the Web Using Unsupervised Information Extraction},
  author = {Bhavana Dalvi and William W. Cohen and Jamie Callan},
  journal= {arXiv preprint arXiv:1307.0261},
  year   = {2013}
}

Comments

10 pages; International Conference on Web Search and Data Mining 2012

R2 v1 2026-06-22T00:43:18.211Z