English

Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability

Computation and Language 2026-07-07 v1 Computers and Society

Abstract

Current AI safety evaluation and benchmarking frameworks predominantly rely on Western-centric culture-agnostic defaults that mask critical regional laws, socio-linguistic nuances, and cultural taboos, leaving Vision-Language Models (VLMs) vulnerable in global deployments. We introduce Pluralis v0.1: a novel multimodal, multi-regional, and multilingual dataset built from a culture-first perspective. Spanning 6,448 prompts across six Asia-Pacific countries (Bangladesh, India, Korea, Pakistan, Singapore, Taiwan) and eight languages, Pluralis diverges from prior work by natively sourcing localized safety hazards rather than adapting Western datasets. Crucially, it introduces a multimodal evaluation paradigm: user text (e.g., "Should I gift this?") and an image referring to "this" (e.g., a clock) - both innocuous in isolation, but synergistically triggering specific legal or cultural violations. Pluralis disentangles universal safety violations from localized cultural appropriateness, establishing the latter as a first-class evaluation axis. To operationalize this, we present Judge-Pluralis, an agreement-gated LLM-as-a-Judge ensemble trained on examples classified in an empirically derived cultural taxonomy. Observing VLM behavior on a subset of the Pluralis surfaces recurring, locale-specific failure modes such as image misidentifications with downstream harm, missed item-context-locale interactions, and inadequate refusals. These failure modes vary systematically across locales and languages, exposing blind spots that globally averaged metrics conceal. Ultimately, Pluralis is not presented as a solved evaluation framework for cultural alignment, but rather as a first step and catalyst for future innovation. We call upon the research community to utilize this foundation to advance the science of multilingual, multicultural evaluation to better support AI cultural alignment globally.

Keywords

Cite

@article{arxiv.2607.06196,
  title  = {Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability},
  author = {Alicia Parrish and Rajat Shinde and Sanket Badhe and Xinyi Bai and Sree Bhargavi Balija and Hua-Rong Chu and Emilio Ferrara and Armstrong Foundjem and Rajat Ghosh and Aakash Gupta and Xuanli He and Ong Chen Hui and Minji Jung and Madhangi Karimanal and Faiza Khan Khattak and Boryoung Kim and Eugenia Kim and Liliya Lavitas and Seok Min Lim and Victor Lu and Jim Moirangthem and Dhivya Nagasubramanian and Deepak Pandita and Sita Rajagopal and Geetha Raju and Evgeniia Razumovskaia and Aravind Reddy and Federico Ricciuti and Nobin Sarwar and Sungpil Shin and Sunayana Sitaram and Snehal Thorat and Tharindu Cyril Weerasooriya and Jasmijn Bastings and Joachim Baumann and Kongtao Chen and Murali Emani and Mariya Hendriksen and Jiho Jin and Jun Seong Kim and Younghoon Ko and Alicja Kwasniewska and Minjae Lee and Tom Wei-cyuan Lin Kashyap Ramanandula Manjusha and Junho Myung and Junyeong Park and Roma Patel and Shyam Ratan and Sudarsun Santhiappan and Priyanka Suresh and Tuesday and Ksheeraj Sai Vepuri Laura Amortegui-Ordonez and Claire Dennis and Minsuk Kahng and Chris Knotz and Alice Oh and Balaraman Ravindran and Soojung Ryu William Bartholomew and Hiwot Tesfaye and Lora Aroyo},
  journal= {arXiv preprint arXiv:2607.06196},
  year   = {2026}
}