English

NorQuAD: Norwegian Question Answering Dataset

Computation and Language 2023-05-04 v1

Abstract

In this paper we present NorQuAD: the first Norwegian question answering dataset for machine reading comprehension. The dataset consists of 4,752 manually created question-answer pairs. We here detail the data collection procedure and present statistics of the dataset. We also benchmark several multilingual and Norwegian monolingual language models on the dataset and compare them against human performance. The dataset will be made freely available.

Cite

@article{arxiv.2305.01957,
  title  = {NorQuAD: Norwegian Question Answering Dataset},
  author = {Sardana Ivanova and Fredrik Aas Andreassen and Matias Jentoft and Sondre Wold and Lilja Øvrelid},
  journal= {arXiv preprint arXiv:2305.01957},
  year   = {2023}
}

Comments

Accepted to NoDaLiDa 2023

R2 v1 2026-06-28T10:24:15.769Z