乌拉尔语识别 (ULI) 2020 共享任务数据集与 Wanca 2017 语料库
计算与语言
2020-08-28 v1
摘要
本文介绍了从互联网爬取的 Wanca 2017 文本语料库,从中收集了用于乌拉尔语识别(ULI)2020 共享任务的稀有乌拉尔语句子。我们描述了 ULI 数据集以及如何使用 Wanca 2017 语料库和来自 Leipzig 语料库集合的不同语言文本构建该数据集。我们还提供了使用 ULI 2020 数据集进行的基线语言识别实验。
引用
@article{arxiv.2008.12169,
title = {Uralic Language Identification (ULI) 2020 shared task dataset and the Wanca 2017 corpus},
author = {Tommi Jauhiainen and Heidi Jauhiainen and Niko Partanen and Krister Lindén},
journal= {arXiv preprint arXiv:2008.12169},
year = {2020}
}