Mainstream large vision-language models (LVLMs) inherently encode cultural biases, highlighting the need for diverse multimodal datasets. To address this gap, we introduce PEARL, a large-scale Arabic multimodal dataset and benchmark explicitly designed for cultural understanding. Constructed through advanced agentic workflows and extensive human-in-the-loop annotations by 37 annotators from across the Arab world, PEARL comprises over 309K multimodal examples spanning ten culturally significant domains covering all Arab countries. We further provide two robust evaluation benchmarks (PEARL and PEARL-LITE) along with a specialized subset (PEARL-X) explicitly developed to assess nuanced cultural variations. Comprehensive evaluations on state-of-the-art open and proprietary LVLMs demonstrate that reasoning-centric instruction alignment substantially improves models' cultural grounding compared to conventional scaling methods. PEARL establishes a foundational resource for advancing culturally-informed multimodal modeling research. All datasets and benchmarks are publicly available.
@article{arxiv.2505.21979,
title = {Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset},
author = {Fakhraddin Alwajih and Samar M. Magdy and Abdellah El Mekki and Omer Nacar and Youssef Nafea and Safaa Taher Abdelfadil and Abdulfattah Mohammed Yahya and Hamzah Luqman and Nada Almarwani and Samah Aloufi and Baraah Qawasmen and Houdaifa Atou and Serry Sibaee and Hamzah A. Alsayadi and Walid Al-Dhabyani and Maged S. Al-shaibani and Aya El Aatar and Nour Qandos and Rahaf Alhamouri and Samar Ahmad and Mohammed Anwar Al-Ghrawi and Aminetou Yacoub and Ruwa AbuHweidi and Vatimetou Mohamed Lemin and Reem Abdel-Salam and Ahlam Bashiti and Aisha Alansari and Ahmed Ashraf and Nora Alturayeif and Alcides Alcoba Inciarte and Adel Ammar and Abdelrahim A. Elmadany and Mohamedou Cheikh Tourad and Ismail Berrada and Mustafa Jarrar and Shady Shehata and Muhammad Abdul-Mageed},
journal= {arXiv preprint arXiv:2505.21979},
year = {2025}
}