English

Improving Reliability of Word Similarity Evaluation by Redesigning Annotation Task and Performance Measure

Computation and Language 2017-02-28 v2

Abstract

We suggest a new method for creating and using gold-standard datasets for word similarity evaluation. Our goal is to improve the reliability of the evaluation, and we do this by redesigning the annotation task to achieve higher inter-rater agreement, and by defining a performance measure which takes the reliability of each annotation decision in the dataset into account.

Keywords

Cite

@article{arxiv.1611.03641,
  title  = {Improving Reliability of Word Similarity Evaluation by Redesigning Annotation Task and Performance Measure},
  author = {Oded Avraham and Yoav Goldberg},
  journal= {arXiv preprint arXiv:1611.03641},
  year   = {2017}
}