English

NaijaRC: A Multi-choice Reading Comprehension Dataset for Nigerian Languages

Computation and Language 2024-05-21 v3

Abstract

In this paper, we create NaijaRC: a new multi-choice Reading Comprehension dataset for three native Nigeria languages that is based on high-school reading comprehension examination. We provide baseline results by performing cross-lingual transfer using existing English RACE and Belebele training dataset based on a pre-trained encoder-only model. Additionally, we provide results by prompting large language models (LLMs) like GPT-4.

Keywords

Cite

@article{arxiv.2308.09768,
  title  = {NaijaRC: A Multi-choice Reading Comprehension Dataset for Nigerian Languages},
  author = {Anuoluwapo Aremu and Jesujoba O. Alabi and Daud Abolade and Nkechinyere F. Aguobi and Shamsuddeen Hassan Muhammad and David Ifeoluwa Adelani},
  journal= {arXiv preprint arXiv:2308.09768},
  year   = {2024}
}

Comments

Accepted to AfricaNLP Workshop at ICLR 2024 (non-archival)