English
Related papers

Related papers: KoBBQ: Korean Bias Benchmark for Question Answerin…

200 papers

This study introduces an innovative multilingual bias evaluation framework for assessing bias in Large Language Models, combining explicit bias assessment through the BBQ benchmark with implicit bias measurement using a prompt-based…

Computers and Society · Computer Science 2025-12-19 Yuxuan Liang , Marwa Mahmoud

As large language models (LLMs) become key advisors in various domains, their cultural sensitivity and reasoning skills are crucial in multicultural environments. We introduce Nunchi-Bench, a benchmark designed to evaluate LLMs' cultural…

Computation and Language · Computer Science 2025-07-08 Kyuhee Kim , Sangah Lee

Recent advances in large audio language models (LALMs) have enabled multilingual speech understanding. However, benchmarks for evaluating LALMs remain scarce for non-English languages, with Korean being one such underexplored case. In this…

Computation and Language · Computer Science 2026-04-23 Jinyoung Kim , Hyeongsoo Lim , Eunseo Seo , Minho Jang , Keunwoo Choi , Seungyoun Shin , Ji Won Yoon

In a highly globalized world, it is important for multi-modal large language models (MLLMs) to recognize and respond correctly to mixed-cultural inputs. For example, a model should correctly identify kimchi (Korean food) in an image both…

Computation and Language · Computer Science 2025-03-24 Jun Seong Kim , Kyaw Ye Thu , Javad Ismayilzada , Junyeong Park , Eunsu Kim , Huzama Ahmad , Na Min An , James Thorne , Alice Oh

As language models are often deployed as chatbot assistants, it becomes a virtue for models to engage in conversations in a user's first language. While these models are trained on a wide range of languages, a comprehensive evaluation of…

Computation and Language · Computer Science 2024-06-18 Seongbo Jang , Seonghyeon Lee , Hwanjo Yu

In enhancing the fairness of Large Language Models (LLMs), evaluating social biases rooted in the cultural contexts of specific linguistic regions is essential. However, most existing Japanese benchmarks heavily rely on translating English…

Computation and Language · Computer Science 2026-04-02 Taihei Shiotani , Masahiro Kaneko , Naoaki Okazaki

An increasing number of studies have examined the social bias of rapidly developed large language models (LLMs). Although most of these studies have focused on bias occurring in a single social attribute, research in social science has…

Computation and Language · Computer Science 2025-07-29 Hitomi Yanaka , Xinqi He , Jie Lu , Namgi Han , Sunjin Oh , Ryoma Kumon , Yuma Matsuoka , Katsuhiko Watabe , Yuko Itatsu

Social biases reflected in language are inherently shaped by cultural norms, which vary significantly across regions and lead to diverse manifestations of stereotypes. Existing evaluations of social bias in large language models (LLMs) for…

Computation and Language · Computer Science 2026-03-26 Taihei Shiotani , Masahiro Kaneko , Ayana Niwa , Yuki Maruyama , Daisuke Oba , Masanari Ohi , Naoaki Okazaki

We present $\textbf{Korean SimpleQA (KoSimpleQA)}$, a benchmark for evaluating factuality in large language models (LLMs) with a focus on Korean cultural knowledge. KoSimpleQA is designed to be challenging yet easy to grade, consisting of…

Computation and Language · Computer Science 2025-10-22 Donghyeon Ko , Yeguk Jin , Kyubyung Chae , Byungwook Lee , Chansong Jo , Sookyo In , Jaehong Lee , Taesup Kim , Donghyun Kwak

For Large Language Models (LLMs) to be effectively deployed in a specific country, they must possess an understanding of the nation's culture and basic knowledge. To this end, we introduce National Alignment, which measures an alignment…

Computation and Language · Computer Science 2024-06-14 Jiyoung Lee , Minwoo Kim , Seungho Kim , Junghwan Kim , Seunghyun Won , Hwaran Lee , Edward Choi

Current social bias benchmarks for Large Language Models (LLMs) primarily rely on predefined question formats like multiple-choice, limiting their ability to reflect the complexity and open-ended nature of real-world interactions. To close…

Computation and Language · Computer Science 2025-10-16 Zhao Liu , Tian Xie , Xueru Zhang

To create culturally inclusive vision-language models (VLMs), developing a benchmark that tests their ability to address culturally relevant questions is essential. Existing approaches typically rely on human annotators, making the process…

Computation and Language · Computer Science 2025-06-02 ChaeHun Park , Yujin Baek , Jaeseok Kim , Yu-Jung Heo , Du-Seong Chang , Jaegul Choo

We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. While prior Korean benchmarks are translated from existing English benchmarks, KMMLU is…

Computation and Language · Computer Science 2024-06-07 Guijin Son , Hanwool Lee , Sungdong Kim , Seungone Kim , Niklas Muennighoff , Taekyoon Choi , Cheonbok Park , Kang Min Yoo , Stella Biderman

Question-answering (QA) and reading comprehension (RC) benchmarks are commonly used for assessing the capabilities of large language models (LLMs) to retrieve and reproduce knowledge. However, we demonstrate that popular QA and RC…

Computation and Language · Computer Science 2026-01-08 Angelie Kraft , Judith Simon , Sonja Schimmler

Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based proxies to detect…

Computation and Language · Computer Science 2026-04-03 Bhaskara Hanuma Vedula , Darshan Anghan , Ishita Goyal , Ponnurangam Kumaraguru , Abhijnan Chakraborty

Stereotype biases in Large Multimodal Models (LMMs) perpetuate harmful societal prejudices, undermining the fairness and equity of AI applications. As LMMs grow increasingly influential, addressing and mitigating inherent biases related to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Vishal Narnaware , Ashmal Vayani , Rohit Gupta , Sirnam Swetha , Mubarak Shah

With the growth of online services, the need for advanced text classification algorithms, such as sentiment analysis and biased text detection, has become increasingly evident. The anonymous nature of online services often leads to the…

Computation and Language · Computer Science 2023-11-14 Dasol Choi , Jooyoung Song , Eunsun Lee , Jinwoo Seo , Heejune Park , Dongbin Na

Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and tasks in languages…

Computation and Language · Computer Science 2024-10-14 Yeeun Kim , Young Rok Choi , Eunkyung Choi , Jinhwan Choi , Hai Jin Park , Wonseok Hwang

This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin American contexts. Despite widespread global deployment, current…

Computers and Society · Computer Science 2025-09-04 Melissa Robles , Catalina Bernal , Denniss Raigoso , Mateo Dulce Rubio

Large Language Models have been shown to demonstrate stereotypical biases in their representations and behavior due to the discriminative nature of the data that they have been trained on. Despite significant progress in the development of…

Computation and Language · Computer Science 2025-10-29 Kaveh Eskandari Miandoab , Mahammed Kamruzzaman , Arshia Gharooni , Gene Louis Kim , Vasanth Sarathy , Ninareh Mehrabi