English

Using Information Content to Evaluate Semantic Similarity in a Taxonomy

cmp-lg 2008-02-03 v1 Computation and Language

Abstract

This paper presents a new measure of semantic similarity in an IS-A taxonomy, based on the notion of information content. Experimental evaluation suggests that the measure performs encouragingly well (a correlation of r = 0.79 with a benchmark set of human similarity judgments, with an upper bound of r = 0.90 for human subjects performing the same task), and significantly better than the traditional edge counting approach (r = 0.66).

Keywords

Cite

@article{arxiv.cmp-lg/9511007,
  title  = {Using Information Content to Evaluate Semantic Similarity in a Taxonomy},
  author = {Philip Resnik},
  journal= {arXiv preprint arXiv:cmp-lg/9511007},
  year   = {2008}
}

Comments

6 pages, 2 postscript figures, uses ijcai95.sty

R2 v1 2026-07-22T09:58:21.380Z