English
Related papers

Related papers: KoBBQ: Korean Bias Benchmark for Question Answerin…

200 papers

With the development of large language models (LLMs), social biases in these LLMs have become a pressing issue. Although there are various benchmarks for social biases across languages, the extent to which Japanese LLMs exhibit social…

Computation and Language · Computer Science 2025-06-16 Hitomi Yanaka , Namgi Han , Ryoma Kumon , Jie Lu , Masashi Takeshita , Ryo Sekizawa , Taisei Kato , Hiromi Arai

Evaluating social biases in language models (LMs) is crucial for ensuring fairness and minimizing the reinforcement of harmful stereotypes in AI systems. Existing benchmarks, such as the Bias Benchmark for Question Answering (BBQ),…

Computation and Language · Computer Science 2025-08-12 Aditya Tomar , Nihar Ranjan Sahoo , Pushpak Bhattacharyya

With the widespread adoption of Large Language Models (LLMs) across various applications, it is empirical to ensure their fairness across all user communities. However, most LLMs are trained and evaluated on Western centric data, with…

Computation and Language · Computer Science 2025-09-30 Abdullah Hashmat , Muhammad Arham Mirza , Agha Ali Raza

Despite the rapid development of large language models (LLMs) for the Korean language, there remains an obvious lack of benchmark datasets that test the requisite Korean cultural and linguistic knowledge. Because many existing Korean…

Computation and Language · Computer Science 2024-07-08 Eunsu Kim , Juyoung Suk , Philhoon Oh , Haneul Yoo , James Thorne , Alice Oh

Measuring social bias in large language models (LLMs) is crucial, but existing bias evaluation methods struggle to assess bias in long-form generation. We propose a Bias Benchmark for Generation (BBG), an adaptation of the Bias Benchmark…

Computation and Language · Computer Science 2025-06-13 Jiho Jin , Woosung Kang , Junho Myung , Alice Oh

It is well documented that NLP models learn social biases, but little work has been done on how these biases manifest in model outputs for applied tasks like question answering (QA). We introduce the Bias Benchmark for QA (BBQ), a dataset…

Computation and Language · Computer Science 2022-03-17 Alicia Parrish , Angelica Chen , Nikita Nangia , Vishakh Padmakumar , Jason Phang , Jana Thompson , Phu Mon Htut , Samuel R. Bowman

Generative large language models (LLMs) have been shown to exhibit harmful biases and stereotypes. While safety fine-tuning typically takes place in English, if at all, these models are being used by speakers of many different languages.…

Computation and Language · Computer Science 2024-07-18 Vera Neplenbroek , Arianna Bisazza , Raquel Fernández

We introduce VoiceBBQ, a spoken extension of the BBQ (Bias Benchmark for Question Answering) - a dataset that measures social bias by presenting ambiguous or disambiguated contexts followed by questions that may elicit stereotypical…

Computation and Language · Computer Science 2025-09-26 Junhyuk Choi , Ro-hoon Oh , Jihwan Seol , Bugeun Kim

Large language models (LLMs) learn not only natural text generation abilities but also social biases against different demographic groups from real-world data. This poses a critical risk when deploying LLM-based applications. Existing…

Computation and Language · Computer Science 2023-05-31 Hwaran Lee , Seokhee Hong , Joonsuk Park , Takyoung Kim , Gunhee Kim , Jung-Woo Ha

Holistically measuring societal biases of large language models is crucial for detecting and reducing ethical risks in highly capable AI models. In this work, we present a Chinese Bias Benchmark dataset that consists of over 100K questions…

Computation and Language · Computer Science 2023-06-29 Yufei Huang , Deyi Xiong

With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an important benchmark format for evaluating stereotypical associations exhibited by…

Computation and Language · Computer Science 2026-04-21 Lance Calvin Lim Gamboa , Yue Feng , Mark Lee

Previous literature has largely shown that Large Language Models (LLMs) perpetuate social biases learnt from their pre-training data. Given the notable lack of resources for social bias evaluation in languages other than English, and for…

Large language models have exhibited significant enhancements in performance across various tasks. However, the complexity of their evaluation increases as these models generate more fluent and coherent content. Current multilingual…

Computation and Language · Computer Science 2024-12-11 Xiaonan Wang , Jinyoung Yeo , Joon-Ho Lim , Hansaem Kim

With the increasing adoption of large language models (LLMs), ensuring their alignment with social norms has become a critical concern. While prior research has examined bias detection in various languages, there remains a significant gap…

Computation and Language · Computer Science 2025-10-23 Farhan Farsi , Shayan Bali , Fatemeh Valeh , Parsa Ghofrani , Alireza Pakniat , Kian Kashfipour , Amir H. Payberah

We introduce KoBALT (Korean Benchmark for Advanced Linguistic Tasks), a comprehensive linguistically-motivated benchmark comprising 700 multiple-choice questions spanning 24 phenomena across five linguistic domains: syntax, semantics,…

Computation and Language · Computer Science 2025-05-23 Hyopil Shin , Sangah Lee , Dongjun Jang , Wooseok Song , Jaeyoon Kim , Chaeyoung Oh , Hyemi Jo , Youngchae Ahn , Sihyun Oh , Hyohyeong Chang , Sunkyoung Kim , Jinsik Lee

Physical commonsense reasoning datasets like PIQA are predominantly English-centric and lack cultural diversity. We introduce Ko-PIQA, a Korean physical commonsense reasoning dataset that incorporates cultural context. Starting from 3.01…

Computation and Language · Computer Science 2025-09-30 Dasol Choi , Jungwhan Kim , Guijin Son

Large Language Models (LLMs) have achieved remarkable success on question answering (QA) tasks, yet they often encode harmful biases that compromise fairness and trustworthiness. Most existing bias mitigation approaches are restricted to…

Computation and Language · Computer Science 2025-09-30 Arti Rani , Shweta Singh , Nihar Ranjan Sahoo , Gaurav Kumar Nayak

Large language models (LLMs) trained on massive corpora demonstrate impressive capabilities in a wide range of tasks. While there are ongoing efforts to adapt these models to languages beyond English, the attention given to their evaluation…

Computation and Language · Computer Science 2024-03-21 Guijin Son , Hanwool Lee , Suwan Kim , Huiseo Kim , Jaecheol Lee , Je Won Yeom , Jihyu Jung , Jung Woo Kim , Songseong Kim

A well-formulated benchmark plays a critical role in spurring advancements in the natural language processing (NLP) field, as it allows objective and precise evaluation of diverse models. As modern language models (LMs) have become more…

Computation and Language · Computer Science 2022-04-12 Dohyeong Kim , Myeongjun Jang , Deuk Sin Kwon , Eric Davis

The rapid advancement of large language models (LLMs) has enabled natural language processing capabilities similar to those of humans, and LLMs are being widely utilized across various societal domains such as education and healthcare.…

Computation and Language · Computer Science 2024-03-19 J. K. Lee , T. M. Chung
‹ Prev 1 2 3 10 Next ›