English

CA-EHN: Commonsense Analogy from E-HowNet

Computation and Language 2020-06-01 v5

Abstract

Embedding commonsense knowledge is crucial for end-to-end models to generalize inference beyond training corpora. However, existing word analogy datasets have tended to be handcrafted, involving permutations of hundreds of words with only dozens of pre-defined relations, mostly morphological relations and named entities. In this work, we model commonsense knowledge down to word-level analogical reasoning by leveraging E-HowNet, an ontology that annotates 88K Chinese words with their structured sense definitions and English translations. We present CA-EHN, the first commonsense word analogy dataset containing 90,505 analogies covering 5,656 words and 763 relations. Experiments show that CA-EHN stands out as a great indicator of how well word representations embed commonsense knowledge. The dataset is publicly available at https://github.com/ckiplab/CA-EHN.

Cite

@article{arxiv.1908.07218,
  title  = {CA-EHN: Commonsense Analogy from E-HowNet},
  author = {Peng-Hsuan Li and Tsan-Yu Yang and Wei-Yun Ma},
  journal= {arXiv preprint arXiv:1908.07218},
  year   = {2020}
}

Comments

In proceedings of LREC 2020

R2 v1 2026-06-23T10:51:52.213Z