中文

SuperSim:瑞典语词语相似度与相关度测试集

计算与语言 2021-04-13 v1

摘要

语言模型 notoriously 难以评估。我们发布SuperSim,一个基于专家人工判断构建的大规模瑞典语相似度与相关度测试集。该测试集由1,360个词对组成,由五名标注者独立判断其相关度与相似度。我们评估了在两个独立瑞典语数据集(即瑞典语Gigaword语料库和瑞典语Wikipedia转储)上训练三种不同模型(Word2Vec、fastText和GloVe),以为未来比较提供基线。我们发布了完全标注的测试集、代码、基线模型与数据。

关键词

引用

@article{arxiv.2104.05228,
  title  = {SuperSim: a test set for word similarity and relatedness in Swedish},
  author = {Simon Hengchen and Nina Tahmasebi},
  journal= {arXiv preprint arXiv:2104.05228},
  year   = {2021}
}

备注

Accepted at NoDaLiDa 2021