引入两个越南语数据集用于评估语义(不)相似性与相关性模型
计算与语言
2018-04-20 v2
摘要
我们为低资源语言越南语提供了两个新颖的数据集,以评估语义相似性模型:ViCon 包含跨词类的同义词和反义词对,从而提供了区分相似性和不相似性的数据。ViSim-400 提供了五种语义关系下的相似度,由人类评判者打分评定。这两个数据集通过标准的共现模型和神经网络模型进行了验证,结果显示其结果与相应的英语数据集相当。
引用
@article{arxiv.1804.05388,
title = {Introducing two Vietnamese Datasets for Evaluating Semantic Models of (Dis-)Similarity and Relatedness},
author = {Kim Anh Nguyen and Sabine Schulte im Walde and Ngoc Thang Vu},
journal= {arXiv preprint arXiv:1804.05388},
year = {2018}
}
备注
The 16th Annual Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL HLT 2018)