SuperSim:瑞典语词语相似度与相关度测试集
计算与语言
2021-04-13 v1
摘要
语言模型 notoriously 难以评估。我们发布SuperSim,一个基于专家人工判断构建的大规模瑞典语相似度与相关度测试集。该测试集由1,360个词对组成,由五名标注者独立判断其相关度与相似度。我们评估了在两个独立瑞典语数据集(即瑞典语Gigaword语料库和瑞典语Wikipedia转储)上训练三种不同模型(Word2Vec、fastText和GloVe),以为未来比较提供基线。我们发布了完全标注的测试集、代码、基线模型与数据。
引用
@article{arxiv.2104.05228,
title = {SuperSim: a test set for word similarity and relatedness in Swedish},
author = {Simon Hengchen and Nina Tahmasebi},
journal= {arXiv preprint arXiv:2104.05228},
year = {2021}
}
备注
Accepted at NoDaLiDa 2021