中文

乌拉尔语识别 (ULI) 2020 共享任务数据集与 Wanca 2017 语料库

计算与语言 2020-08-28 v1

摘要

本文介绍了从互联网爬取的 Wanca 2017 文本语料库,从中收集了用于乌拉尔语识别(ULI)2020 共享任务的稀有乌拉尔语句子。我们描述了 ULI 数据集以及如何使用 Wanca 2017 语料库和来自 Leipzig 语料库集合的不同语言文本构建该数据集。我们还提供了使用 ULI 2020 数据集进行的基线语言识别实验。

关键词

引用

@article{arxiv.2008.12169,
  title  = {Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpus},
  author = {Tommi Jauhiainen and Heidi Jauhiainen and Niko Partanen and Krister Lindén},
  journal= {arXiv preprint arXiv:2008.12169},
  year   = {2020}
}