中文

引入两个越南语数据集用于评估语义(不)相似性与相关性模型

计算与语言 2018-04-20 v2

摘要

我们为低资源语言越南语提供了两个新颖的数据集,以评估语义相似性模型:ViCon 包含跨词类的同义词和反义词对,从而提供了区分相似性和不相似性的数据。ViSim-400 提供了五种语义关系下的相似度,由人类评判者打分评定。这两个数据集通过标准的共现模型和神经网络模型进行了验证,结果显示其结果与相应的英语数据集相当。

关键词

引用

@article{arxiv.1804.05388,
  title  = {Introducing two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness},
  author = {Kim Anh Nguyen and Sabine Schulte im Walde and Ngoc Thang Vu},
  journal= {arXiv preprint arXiv:1804.05388},
  year   = {2018}
}

备注

The 16th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT 2018)