English

HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language

Computation and Language 2023-05-30 v1

Abstract

This paper presents HaVQA, the first multimodal dataset for visual question-answering (VQA) tasks in the Hausa language. The dataset was created by manually translating 6,022 English question-answer pairs, which are associated with 1,555 unique images from the Visual Genome dataset. As a result, the dataset provides 12,044 gold standard English-Hausa parallel sentences that were translated in a fashion that guarantees their semantic match with the corresponding visual information. We conducted several baseline experiments on the dataset, including visual question answering, visual question elicitation, text-only and multimodal machine translation.

Keywords

Cite

@article{arxiv.2305.17690,
  title  = {HaVQA: A Dataset for Visual Question Answering and Multimodal Research in Hausa Language},
  author = {Shantipriya Parida and Idris Abdulmumin and Shamsuddeen Hassan Muhammad and Aneesh Bose and Guneet Singh Kohli and Ibrahim Said Ahmad and Ketan Kotwal and Sayan Deb Sarkar and Ondřej Bojar and Habeebah Adamu Kakudi},
  journal= {arXiv preprint arXiv:2305.17690},
  year   = {2023}
}

Comments

Accepted at ACL 2023 as a long paper (Findings)