中文
相关论文

相关论文: BharatBBQ: A Multilingual Bias Benchmark for Quest…

200 篇论文

With the widespread adoption of Large Language Models (LLMs) across various applications, it is empirical to ensure their fairness across all user communities. However, most LLMs are trained and evaluated on Western centric data, with…

计算与语言 · 计算机科学 2025-09-30 Abdullah Hashmat , Muhammad Arham Mirza , Agha Ali Raza

Generative large language models (LLMs) have been shown to exhibit harmful biases and stereotypes. While safety fine-tuning typically takes place in English, if at all, these models are being used by speakers of many different languages.…

计算与语言 · 计算机科学 2024-07-18 Vera Neplenbroek , Arianna Bisazza , Raquel Fernández

The pervasive influence of social biases in language data has sparked the need for benchmark datasets that capture and evaluate these biases in Large Language Models (LLMs). Existing efforts predominantly focus on English language and the…

The Bias Benchmark for Question Answering (BBQ) is designed to evaluate social biases of language models (LMs), but it is not simple to adapt this benchmark to cultural contexts other than the US because social biases depend heavily on the…

计算与语言 · 计算机科学 2024-01-26 Jiho Jin , Jiseon Kim , Nayeon Lee , Haneul Yoo , Alice Oh , Hwaran Lee

With the development of large language models (LLMs), social biases in these LLMs have become a pressing issue. Although there are various benchmarks for social biases across languages, the extent to which Japanese LLMs exhibit social…

计算与语言 · 计算机科学 2025-06-16 Hitomi Yanaka , Namgi Han , Ryoma Kumon , Jie Lu , Masashi Takeshita , Ryo Sekizawa , Taisei Kato , Hiromi Arai

It is well documented that NLP models learn social biases, but little work has been done on how these biases manifest in model outputs for applied tasks like question answering (QA). We introduce the Bias Benchmark for QA (BBQ), a dataset…

Large Language Models (LLMs) perform well on unseen tasks in English, but their abilities in non English languages are less explored due to limited benchmarks and training data. To bridge this gap, we introduce the Indic QA Benchmark, a…

Previous literature has largely shown that Large Language Models (LLMs) perpetuate social biases learnt from their pre-training data. Given the notable lack of resources for social bias evaluation in languages other than English, and for…

We introduce VoiceBBQ, a spoken extension of the BBQ (Bias Benchmark for Question Answering) - a dataset that measures social bias by presenting ambiguous or disambiguated contexts followed by questions that may elicit stereotypical…

计算与语言 · 计算机科学 2025-09-26 Junhyuk Choi , Ro-hoon Oh , Jihwan Seol , Bugeun Kim

This study introduces an innovative multilingual bias evaluation framework for assessing bias in Large Language Models, combining explicit bias assessment through the BBQ benchmark with implicit bias measurement using a prompt-based…

计算机与社会 · 计算机科学 2025-12-19 Yuxuan Liang , Marwa Mahmoud

With the increasing adoption of large language models (LLMs), ensuring their alignment with social norms has become a critical concern. While prior research has examined bias detection in various languages, there remains a significant gap…

计算与语言 · 计算机科学 2025-10-23 Farhan Farsi , Shayan Bali , Fatemeh Valeh , Parsa Ghofrani , Alireza Pakniat , Kian Kashfipour , Amir H. Payberah

With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an important benchmark format for evaluating stereotypical associations exhibited by…

计算与语言 · 计算机科学 2026-04-21 Lance Calvin Lim Gamboa , Yue Feng , Mark Lee

Measuring social bias in large language models (LLMs) is crucial, but existing bias evaluation methods struggle to assess bias in long-form generation. We propose a Bias Benchmark for Generation (BBG), an adaptation of the Bias Benchmark…

计算与语言 · 计算机科学 2025-06-13 Jiho Jin , Woosung Kang , Junho Myung , Alice Oh

This paper addresses the critical gap in evaluating bias in multilingual Large Language Models (LLMs), with a specific focus on Spanish language within culturally-aware Latin American contexts. Despite widespread global deployment, current…

计算机与社会 · 计算机科学 2025-09-04 Melissa Robles , Catalina Bernal , Denniss Raigoso , Mateo Dulce Rubio

Large Language Models have been shown to demonstrate stereotypical biases in their representations and behavior due to the discriminative nature of the data that they have been trained on. Despite significant progress in the development of…

Existing studies on fairness are largely Western-focused, making them inadequate for culturally diverse countries such as India. To address this gap, we introduce INDIC-BIAS, a comprehensive India-centric benchmark designed to evaluate…

计算与语言 · 计算机科学 2025-07-01 Janki Atul Nawale , Mohammed Safi Ur Rahman Khan , Janani D , Mansi Gupta , Danish Pruthi , Mitesh M. Khapra

Stereotype biases in Large Multimodal Models (LMMs) perpetuate harmful societal prejudices, undermining the fairness and equity of AI applications. As LMMs grow increasingly influential, addressing and mitigating inherent biases related to…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Vishal Narnaware , Ashmal Vayani , Rohit Gupta , Sirnam Swetha , Mubarak Shah

Large Language Models (LLMs), now used daily by millions, can encode societal biases, exposing their users to representational harms. A large body of scholarship on LLM bias exists but it predominantly adopts a Western-centric frame and…

计算与语言 · 计算机科学 2024-08-12 Khyati Khandelwal , Manuel Tonneau , Andrew M. Bean , Hannah Rose Kirk , Scott A. Hale

Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit biases when identity is conveyed indirectly. Existing benchmarks use name based proxies to detect…

计算与语言 · 计算机科学 2026-04-03 Bhaskara Hanuma Vedula , Darshan Anghan , Ishita Goyal , Ponnurangam Kumaraguru , Abhijnan Chakraborty

Large Language Models (LLMs) have achieved significant success in recent years; yet, issues of intrinsic gender bias persist, especially in non-English languages. Although current research mostly emphasizes English, the linguistic and…

‹ 上一页 1 2 3 10 下一页 ›