DaLAJ——瑞典语语言可接受性判断数据集:格式、基线、共享
摘要
我们提出 DaLAJ 1.0,一个瑞典语语言可接受性判断数据集(Dataset for Linguistic Acceptability Judgments for Swedish),其首个版本包含 9 596 个句子;并介绍了使用该数据集进行二分类任务的初始实验。DaLAJ 基于 SweLL 二语学习者数据,该数据由不同熟练程度的作文组成。为确保数据集在 GDPR 法规下可自由获取,我们对学习者作文进行了句子打乱处理,并移除了部分关于学习者的元数据,仅为每个句子保留母语信息及作文所属课程级别。我们以规范化后的学习者语言作为 DaLAJ 句子的基础,且每句仅保留一个错误。对于句中使用的每个修正标签,我们重复相同的句子。DaLAJ 1.0 使用了四类错误(SweLL 中共有 35 类可用),均与词汇或构词选择相关。我们使用 BERT embeddings 在 DaLAJ 1.0 上的二分类基线准确率为 58%。该数据集已纳入 SwedishGlue(瑞典语 SuperLim)基准。下文我们描述了数据集格式、首次实验、我们的见解以及所选数据共享方法的动机。
引用
@article{arxiv.2105.06681,
title = {DaLAJ - a dataset for linguistic acceptability judgments for Swedish: Format, baseline, sharing},
author = {Elena Volodina and Yousuf Ali Mohammed and Julia Klezl},
journal= {arXiv preprint arXiv:2105.06681},
year = {2021}
}
备注
This is an extended version of an article accepted to the 10th NLP4CALL workshop (2021), Link\"oping Electronic Conference Proceedings 177, ISSN: 1650-3740 (online). In the extended version (available at arXiv) we have added a description of an experiment and baseline results to the dataset description accepted for NLP4CALL publication