English

Intrinsic Subspace Evaluation of Word Embedding Representations

Computation and Language 2016-06-28 v1

Abstract

We introduce a new methodology for intrinsic evaluation of word representations. Specifically, we identify four fundamental criteria based on the characteristics of natural language that pose difficulties to NLP systems; and develop tests that directly show whether or not representations contain the subspaces necessary to satisfy these criteria. Current intrinsic evaluations are mostly based on the overall similarity or full-space similarity of words and thus view vector representations as points. We show the limits of these point-based intrinsic evaluations. We apply our evaluation methodology to the comparison of a count vector model and several neural network models and demonstrate important properties of these models.

Keywords

Cite

@article{arxiv.1606.07902,
  title  = {Intrinsic Subspace Evaluation of Word Embedding Representations},
  author = {Yadollah Yaghoobzadeh and Hinrich Schütze},
  journal= {arXiv preprint arXiv:1606.07902},
  year   = {2016}
}

Comments

Long paper accepted in ACL2016

R2 v1 2026-06-22T14:34:07.651Z