English

NativQA: Multilingual Culturally-Aligned Natural Query for LLMs

Computation and Language 2025-06-02 v3 Artificial Intelligence

Abstract

Natural Question Answering (QA) datasets play a crucial role in evaluating the capabilities of large language models (LLMs), ensuring their effectiveness in real-world applications. Despite the numerous QA datasets that have been developed and some work has been done in parallel, there is a notable lack of a framework and large scale region-specific datasets queried by native users in their own languages. This gap hinders the effective benchmarking and the development of fine-tuned models for regional and cultural specificities. In this study, we propose a scalable, language-independent framework, NativQA, to seamlessly construct culturally and regionally aligned QA datasets in native languages, for LLM evaluation and tuning. We demonstrate the efficacy of the proposed framework by designing a multilingual natural QA dataset, MultiNativQA, consisting of ~64k manually annotated QA pairs in seven languages, ranging from high to extremely low resource, based on queries from native speakers from 9 regions covering 18 topics. We benchmark open- and closed-source LLMs with the MultiNativQA dataset. We made the MultiNativQA dataset(https://huggingface.co/datasets/QCRI/MultiNativQA), and other experimental scripts(https://gitlab.com/nativqa/multinativqa) publicly available for the community.

Keywords

Cite

@article{arxiv.2407.09823,
  title  = {NativQA: Multilingual Culturally-Aligned Natural Query for LLMs},
  author = {Md. Arid Hasan and Maram Hasanain and Fatema Ahmad and Sahinur Rahman Laskar and Sunaya Upadhyay and Vrunda N Sukhadia and Mucahid Kutlu and Shammur Absar Chowdhury and Firoj Alam},
  journal= {arXiv preprint arXiv:2407.09823},
  year   = {2025}
}

Comments

LLMs, Native, Multilingual, Language Diversity, Contextual Understanding, Minority Languages, Culturally Informed, Foundation Models, Large Language Models

R2 v1 2026-06-28T17:39:37.303Z