English

ArabicaQA: A Comprehensive Dataset for Arabic Question Answering

Computation and Language 2024-03-27 v1 Information Retrieval

Abstract

In this paper, we address the significant gap in Arabic natural language processing (NLP) resources by introducing ArabicaQA, the first large-scale dataset for machine reading comprehension and open-domain question answering in Arabic. This comprehensive dataset, consisting of 89,095 answerable and 3,701 unanswerable questions created by crowdworkers to look similar to answerable ones, along with additional labels of open-domain questions marks a crucial advancement in Arabic NLP resources. We also present AraDPR, the first dense passage retrieval model trained on the Arabic Wikipedia corpus, specifically designed to tackle the unique challenges of Arabic text retrieval. Furthermore, our study includes extensive benchmarking of large language models (LLMs) for Arabic question answering, critically evaluating their performance in the Arabic language context. In conclusion, ArabicaQA, AraDPR, and the benchmarking of LLMs in Arabic question answering offer significant advancements in the field of Arabic NLP. The dataset and code are publicly accessible for further research https://github.com/DataScienceUIBK/ArabicaQA.

Keywords

Cite

@article{arxiv.2403.17848,
  title  = {ArabicaQA: A Comprehensive Dataset for Arabic Question Answering},
  author = {Abdelrahman Abdallah and Mahmoud Kasem and Mahmoud Abdalla and Mohamed Mahmoud and Mohamed Elkasaby and Yasser Elbendary and Adam Jatowt},
  journal= {arXiv preprint arXiv:2403.17848},
  year   = {2024}
}

Comments

Accepted at SIGIR 2024

R2 v1 2026-06-28T15:34:23.895Z