English
Related papers

Related papers: PakBBQ: A Culturally Adapted Bias Benchmark for QA

200 papers

In this work, we introduce BLUCK, a new dataset designed to measure the performance of Large Language Models (LLMs) in Bengali linguistic understanding and cultural knowledge. Our dataset comprises 2366 multiple-choice questions (MCQs)…

Computation and Language · Computer Science 2026-01-21 Daeen Kabir , Minhajur Rahman Chowdhury Mahim , Sheikh Shafayat , Adnan Sadik , Arian Ahmed , Eunsu Kim , Alice Oh

The rapid adoption of Large Language Models (LLMs) has raised important concerns about the factual reliability of their outputs, particularly in low-resource languages such as Urdu. Existing automated fact-checking systems are predominantly…

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning…

Computation and Language · Computer Science 2025-10-29 Tyler A. Chang , Catherine Arnett , Abdelrahman Eldesokey , Abdelrahman Sadallah , Abeer Kashar , Abolade Daud , Abosede Grace Olanihun , Adamu Labaran Mohammed , Adeyemi Praise , Adhikarinayum Meerajita Sharma , Aditi Gupta , Afitab Iyigun , Afonso Simplício , Ahmed Essouaied , Aicha Chorana , Akhil Eppa , Akintunde Oladipo , Akshay Ramesh , Aleksei Dorkin , Alfred Malengo Kondoro , Alham Fikri Aji , Ali Eren Çetintaş , Allan Hanbury , Alou Dembele , Alp Niksarli , Álvaro Arroyo , Amin Bajand , Amol Khanna , Ana Chkhaidze , Ana Condez , Andiswa Mkhonto , Andrew Hoblitzell , Andrew Tran , Angelos Poulis , Anirban Majumder , Anna Vacalopoulou , Annette Kuuipolani Kanahele Wong , Annika Simonsen , Anton Kovalev , Ashvanth. S , Ayodeji Joseph Lana , Barkin Kinay , Bashar Alhafni , Benedict Cibalinda Busole , Bernard Ghanem , Bharti Nathani , Biljana Stojanovska Đurić , Bola Agbonile , Bragi Bergsson , Bruce Torres Fischer , Burak Tutar , Burcu Alakuş Çınar , Cade J. Kanoniakapueo Kane , Can Udomcharoenchaikit , Catherine Arnett , Chadi Helwe , Chaithra Reddy Nerella , Chen Cecilia Liu , Chiamaka Glory Nwokolo , Cristina España-Bonet , Cynthia Amol , DaeYeop Lee , Dana Arad , Daniil Dzenhaliou , Daria Pugacheva , Dasol Choi , Daud Abolade , David Liu , David Semedo , Deborah Popoola , Deividas Mataciunas , Delphine Nyaboke , Dhyuthy Krishna Kumar , Diogo Glória-Silva , Diogo Tavares , Divyanshu Goyal , DongGeon Lee , Ebele Nwamaka Anajemba , Egonu Ngozi Grace , Elena Mickel , Elena Tutubalina , Elias Herranen , Emile Anand , Emmanuel Habumuremyi , Emuobonuvie Maria Ajiboye , Eryawan Presma Yulianrifat , Esther Adenuga , Ewa Rudnicka , Faith Olabisi Itiola , Faran Taimoor Butt , Fathima Thekkekara , Fatima Haouari , Filbert Aurelian Tjiaranata , Firas Laakom , Francesca Grasso , Francesco Orabona , Francesco Periti , Gbenga Kayode Solomon , Gia Nghia Ngo , Gloria Udhehdhe-oze , Gonçalo Martins , Gopi Naga Sai Ram Challagolla , Guijin Son , Gulnaz Abdykadyrova , Hafsteinn Einarsson , Hai Hu , Hamidreza Saffari , Hamza Zaidi , Haopeng Zhang , Harethah Abu Shairah , Harry Vuong , Hele-Andra Kuulmets , Houda Bouamor , Hwanjo Yu , Iben Nyholm Debess , İbrahim Ethem Deveci , Ikhlasul Akmal Hanif , Ikhyun Cho , Inês Calvo , Inês Vieira , Isaac Manzi , Ismail Daud , Itay Itzhak , Iuliia , Alekseenko , Ivan Belashkin , Ivan Spada , Ivan Zhelyazkov , Jacob Brinton , Jafar Isbarov , Jaka Čibej , Jan Čuhel , Jan Kocoń , Jauza Akbar Krito , Jebish Purbey , Jennifer Mickel , Jennifer Za , Jenny Kunz , Jihae Jeong , Jimena Tena Dávalos , Jinu Lee , João Magalhães , John Yi , Jongin Kim , Joseph Chataignon , Joseph Marvin Imperial , Jubeerathan Thevakumar , Judith Land , Junchen Jiang , Jungwhan Kim , Kairit Sirts , Kamesh R , Kamesh V , Kanda Patrick Tshinu , Kätriin Kukk , Kaustubh Ponkshe , Kavsar Huseynova , Ke He , Kelly Buchanan , Kengatharaiyer Sarveswaran , Kerem Zaman , Khalil Mrini , Kian Kyars , Krister Kruusmaa , Kusum Chouhan , Lainitha Krishnakumar , Laura Castro Sánchez , Laura Porrino Moscoso , Leshem Choshen , Levent Sencan , Lilja Øvrelid , Lisa Alazraki , Lovina Ehimen-Ugbede , Luheerathan Thevakumar , Luxshan Thavarasa , Mahnoor Malik , Mamadou K. Keita , Mansi Jangid , Marco De Santis , Marcos García , Marek Suppa , Mariam D'Ciofalo , Marii Ojastu , Maryam Sikander , Mausami Narayan , Maximos Skandalis , Mehak Mehak , Mehmet İlteriş Bozkurt , Melaku Bayu Workie , Menan Velayuthan , Michael Leventhal , Michał Marcińczuk , Mirna Potočnjak , Mohammadamin Shafiei , Mridul Sharma , Mrityunjaya Indoria , Muhammad Ravi Shulthan Habibi , Murat Kolić , Nada Galant , Naphat Permpredanun , Narada Maugin , Nicholas Kluge Corrêa , Nikola Ljubešić , Nirmal Thomas , Nisansa de Silva , Nisheeth Joshi , Nitish Ponkshe , Nizar Habash , Nneoma C. Udeze , Noel Thomas , Noémi Ligeti-Nagy , Nouhoum Coulibaly , Nsengiyumva Faustin , Odunayo Kareemat Buliaminu , Odunayo Ogundepo , Oghojafor Godswill Fejiro , Ogundipe Blessing Funmilola , Okechukwu God'spraise , Olanrewaju Samuel , Olaoye Deborah Oluwaseun , Olasoji Akindejoye , Olga Popova , Olga Snissarenko , Onyinye Anulika Chiemezie , Orkun Kinay , Osman Tursun , Owoeye Tobiloba Moses , Oyelade Oluwafemi Joshua , Oyesanmi Fiyinfoluwa , Pablo Gamallo , Pablo Rodríguez Fernández , Palak Arora , Pedro Valente , Peter Rupnik , Philip Oghenesuowho Ekiugbo , Pramit Sahoo , Prokopis Prokopidis , Pua Niau-Puhipau , Quadri Yahya , Rachele Mignone , Raghav Singhal , Ram Mohan Rao Kadiyala , Raphael Merx , Rapheal Afolayan , Ratnavel Rajalakshmi , Rishav Ghosh , Romina Oji , Ron Kekeha Solis , Rui Guerra , Rushikesh Zawar , Sa'ad Nasir Bashir , Saeed Alzaabi , Sahil Sandeep , Sai Pavan Batchu , SaiSandeep Kantareddy , Salsabila Zahirah Pranida , Sam Buchanan , Samuel Rutunda , Sander Land , Sarah Sulollari , Sardar Ali , Saroj Sapkota , Saulius Tautvaisas , Sayambhu Sen , Sayantani Banerjee , Sebastien Diarra , SenthilNathan. M , Sewoong Lee , Shaan Shah , Shankar Venkitachalam , Sharifa Djurabaeva , Sharon Ibejih , Shivanya Shomir Dutta , Siddhant Gupta , Silvia Paniagua Suárez , Sina Ahmadi , Sivasuthan Sukumar , Siyuan Song , Snegha A. , Sokratis Sofianopoulos , Sona Elza Simon , Sonja Benčina , Sophie Gvasalia , Sphurti Kirit More , Spyros Dragazis , Stephan P. Kaufhold , Suba. S , Sultan AlRashed , Surangika Ranathunga , Taiga Someya , Taja Kuzman Pungeršek , Tal Haklay , Tasi'u Jibril , Tatsuya Aoyama , Tea Abashidze , Terenz Jomar Dela Cruz , Terra Blevins , Themistoklis Nikas , Theresa Dora Idoko , Thu Mai Do , Tilek Chubakov , Tommaso Gargiani , Uma Rathore , Uni Johannesen , Uwuma Doris Ugwu , Vallerie Alexandra Putra , Vanya Bannihatti Kumar , Varsha Jeyarajalingam , Varvara Arzt , Vasudevan Nedumpozhimana , Viktoria Ondrejova , Viktoryia Horbik , Vishnu Vardhan Reddy Kummitha , Vuk Dinić , Walelign Tewabe Sewunetie , Winston Wu , Xiaojing Zhao , Yacouba Diarra , Yaniv Nikankin , Yash Mathur , Yixi Chen , Yiyuan Li , Yolanda Xavier , Yonatan Belinkov , Yusuf Ismail Abayomi , Zaid Alyafeai , Zhengyang Shan , Zhi Rui Tam , Zilu Tang , Zuzana Nadova , Baber Abbasi , Stella Biderman , David Stap , Duygu Ataman , Fabian Schmidt , Hila Gonen , Jiayi Wang , David Ifeoluwa Adelani

As language models (LMs) become increasingly powerful and widely used, it is important to quantify them for sociodemographic bias with potential for harm. Prior measures of bias are sensitive to perturbations in the templates designed to…

Computation and Language · Computer Science 2024-08-09 Vipul Gupta , Pranav Narayanan Venkit , Hugo Laurençon , Shomir Wilson , Rebecca J. Passonneau

Large Language Models are increasingly applied in the petroleum industry, highlighting the need for a domain-specific evaluation framework. This study develops a benchmark for LLMs in petroleum engineering, including a three-stage process…

Artificial Intelligence · Computer Science 2026-05-28 Xiang Wang , Tingting Zhang , Sen Wang , Ying Wu , Heng Meng , Peng Zhou , Peng Li

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support…

This paper introduces a comprehensive benchmark for evaluating how Large Language Models (LLMs) respond to linguistic shibboleths: subtle linguistic markers that can inadvertently reveal demographic attributes such as gender, social class,…

Computation and Language · Computer Science 2025-08-08 Julia Kharchenko , Tanya Roosta , Aman Chadha , Chirag Shah

Large Multimodal Models (LMMs) are typically trained on vast corpora of image-text data but are often limited in linguistic coverage, leading to biased and unfair outputs across languages. While prior work has explored multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Ananya Raval , Aravind Narayanan , Vahid Reza Khazaie , Shaina Raza

This paper presents a comprehensive evaluation framework for aligning Persian Large Language Models (LLMs) with critical ethical dimensions, including safety, fairness, and social norms. It addresses the gaps in existing LLM evaluation…

Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine,…

Arabic is one of the most widely spoken languages in the world, yet efforts to develop and evaluate Large Language Models (LLMs) for Arabic remain relatively limited. Most existing Arabic benchmarks focus on linguistic, cultural, or…

Large Language Models (LLMs) are increasingly deployed in resume screening pipelines. Although explicit PII (e.g., names) is commonly redacted, resumes typically retain subtle sociocultural markers (languages, co-curricular activities,…

Computers and Society · Computer Science 2026-05-06 Bryan Chen Zhengyu Tan , Shaun Khoo , Bich Ngoc Doan , Zhengyuan Liu , Nancy F. Chen , Roy Ka-Wei Lee

Large Language Models (LLMs) often exhibit cultural biases due to training data dominated by high-resource languages like English and Chinese. This poses challenges for accurately representing and evaluating diverse cultural contexts,…

Computation and Language · Computer Science 2025-08-11 Zhong Ken Hew , Jia Xin Low , Sze Jue Yang , Chee Seng Chan

Current evaluation benchmarks for question answering (QA) in Indic languages often rely on machine translation of existing English datasets. This approach suffers from bias and inaccuracies inherent in machine translation, leading to…

Computation and Language · Computer Science 2024-05-01 Vaishak Narayanan , Prabin Raj KP , Saifudheen Nouphal

Predictive analysis is a cornerstone of modern decision-making, with applications in various domains. Large Language Models (LLMs) have emerged as powerful tools in enabling nuanced, knowledge-intensive conversations, thus aiding in complex…

Computation and Language · Computer Science 2025-05-26 Qin Chen , Yuanyi Ren , Xiaojun Ma , Yuyang Shi

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing…

As large language models (LLMs) are increasingly deployed across diverse linguistic and cultural contexts, understanding their behavior in both factual and disputable scenarios is essential, especially when their outputs may shape public…

Computation and Language · Computer Science 2025-06-30 Sean Kim , Hyuhng Joon Kim

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce \textbf{\textit{CultSportQA}}, a benchmark designed to assess LMs'…

Cultural Intelligence (CQ) refers to the ability to understand unfamiliar cultural contexts, a crucial skill for large language models (LLMs) to effectively engage with globally diverse users. Existing studies often focus on explicitly…

Computation and Language · Computer Science 2025-10-10 Ziyi Liu , Priyanka Dey , Jen-tse Huang , Zhenyu Zhao , Bowen Jiang , Rahul Gupta , Yang Liu , Yao Du , Jieyu Zhao