English
Related papers

Related papers: On the Unintended Social Bias of Training Language…

200 papers

Despite the huge progress in myriad generation tasks, pretrained language models (LMs) such as GPT2 still tend to generate repetitive texts with maximization-based decoding algorithms for open-ended generation. We attribute their…

Computation and Language · Computer Science 2023-07-06 Jian Guan , Minlie Huang

Technology for language generation has advanced rapidly, spurred by advancements in pre-training large models on massive amounts of data and the need for intelligent agents to communicate in a natural manner. While techniques can…

Computation and Language · Computer Science 2021-06-24 Emily Sheng , Kai-Wei Chang , Premkumar Natarajan , Nanyun Peng

Language serves as a powerful tool for the manifestation of societal belief systems. In doing so, it also perpetuates the prevalent biases in our society. Gender bias is one of the most pervasive biases in our society and is seen in online…

Computation and Language · Computer Science 2023-10-27 Rishav Hada , Agrima Seth , Harshita Diddee , Kalika Bali

Pre-trained language models learn socially harmful biases from their training corpora, and may repeat these biases when used for generation. We study gender biases associated with the protagonist in model-generated stories. Such biases may…

Computation and Language · Computer Science 2021-09-15 Tenghao Huang , Faeze Brahman , Vered Shwartz , Snigdha Chaturvedi

Gender bias exists in natural language datasets which neural language models tend to learn, resulting in biased text generation. In this research, we propose a debiasing approach based on the loss function modification. We introduce a new…

Computation and Language · Computer Science 2019-06-05 Yusu Qian , Urwa Muaz , Ben Zhang , Jae Won Hyun

Text-to-Image (T2I) models have transformed visual content creation, producing highly realistic images from natural language prompts. However, concerns persist around their potential to replicate and magnify existing societal biases. To…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Sedat Porikli , Vedat Porikli

Pretrained machine learning models are known to perpetuate and even amplify existing biases in data, which can result in unfair outcomes that ultimately impact user experience. Therefore, it is crucial to understand the mechanisms behind…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Laura Cabello , Emanuele Bugliarello , Stephanie Brandl , Desmond Elliott

Pre-trained language models (PLMs) are trained on data that inherently contains gender biases, leading to undesirable impacts. Traditional debiasing methods often rely on external corpora, which may lack quality, diversity, or demographic…

Computation and Language · Computer Science 2025-03-13 Liu Yu , Ludie Guo , Ping Kuang , Fan Zhou

We examine whether neural natural language processing (NLP) systems reflect historical biases in training data. We define a general benchmark to quantify gender bias in a variety of neural NLP tasks. Our empirical evaluation with…

Computation and Language · Computer Science 2019-06-03 Kaiji Lu , Piotr Mardziel , Fangjing Wu , Preetam Amancharla , Anupam Datta

Image captioning models are known to perpetuate and amplify harmful societal bias in the training set. In this work, we aim to mitigate such gender bias in image captioning models. While prior work has addressed this problem by forcing…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Yusuke Hirota , Yuta Nakashima , Noa Garcia

The representations in large language models contain multiple types of gender information. We focus on two types of such signals in English texts: factual gender information, which is a grammatical or semantic property, and gender bias,…

Computation and Language · Computer Science 2022-06-23 Tomasz Limisiewicz , David Mareček

AI-based systems such as language models have been shown to replicate and even amplify social biases reflected in their training data. Among other questionable behaviors, this can lead to AI-generated text--and text suggestions--that…

Computation and Language · Computer Science 2026-02-19 Connor Baumler , Hal Daumé

The blind application of machine learning runs the risk of amplifying biases present in data. Such a danger is facing us with word embedding, a popular framework to represent text data as vectors which has been used in many machine learning…

Computation and Language · Computer Science 2016-07-25 Tolga Bolukbasi , Kai-Wei Chang , James Zou , Venkatesh Saligrama , Adam Kalai

Mitigating biases in generative AI and, particularly in text-to-image models, is of high importance given their growing implications in society. The biased datasets used for training pose challenges in ensuring the responsible development…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Carolina Lopez Olmos , Alexandros Neophytou , Sunando Sengupta , Dim P. Papadopoulos

Gender bias in pretrained language models (PLMs) poses significant social and ethical challenges. Despite growing awareness, there is a lack of comprehensive investigation into how different models internally represent and propagate such…

Computation and Language · Computer Science 2025-03-11 Mahdi Zakizadeh , Mohammad Taher Pilehvar

Language models have been shown to propagate social bias through their output, particularly in the representation of gender and ethnicity. This paper investigates gender and ethnicity biases in AI-generated occupational stories.…

Computation and Language · Computer Science 2025-09-08 Martha O. Dimgba , Sharon Oba , Ameeta Agrawal , Philippe J. Giabbanelli

The increasing use of Large Language Models (LLMs) in a large variety of domains has sparked worries about how easily they can perpetuate stereotypes and contribute to the generation of biased content. With a focus on gender and…

Computation and Language · Computer Science 2025-07-28 Gioele Giachino , Marco Rondina , Antonio Vetrò , Riccardo Coppola , Juan Carlos De Martin

Neural machine translation has significantly pushed forward the quality of the field. However, there are remaining big issues with the output translations and one of them is fairness. Neural models are trained on large text corpora which…

Computation and Language · Computer Science 2019-06-04 Joel Escudé Font , Marta R. Costa-jussà

After just a few hundred training updates, a standard probabilistic model for language generation has likely not yet learnt many semantic or syntactic rules of natural language, making it difficult to estimate the probability distribution…

Computation and Language · Computer Science 2023-06-26 Clara Meister , Wojciech Stokowiec , Tiago Pimentel , Lei Yu , Laura Rimell , Adhiguna Kuncoro

Many text corpora exhibit socially problematic biases, which can be propagated or amplified in the models trained on such data. For example, doctor cooccurs more frequently with male pronouns than female pronouns. In this study we (i)…

Computation and Language · Computer Science 2019-04-08 Shikha Bordia , Samuel R. Bowman