English
Related papers

Related papers: Augmenting Bias Detection in LLMs Using Topologica…

200 papers

With the impressive performance in various downstream tasks, large language models (LLMs) have been widely integrated into production pipelines, like recruitment and recommendation systems. A known issue of models trained on natural…

Computation and Language · Computer Science 2025-01-22 Damin Zhang , Yi Zhang , Geetanjali Bihani , Julia Rayz

Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application in critical settings. However, the majority of existing methodologies for identifying biases in…

Computation and Language · Computer Science 2025-02-04 Erica Coppolillo , Giuseppe Manco , Luca Maria Aiello

In machine learning, a bias occurs whenever training sets are not representative for the test data, which results in unreliable models. The most common biases in data are arguably class imbalance and covariate shift. In this work, we aim to…

Machine Learning · Computer Science 2018-04-04 Patrick Glauner , Radu State , Petko Valtchev , Diogo Duarte

Large Language Models have garnered significant attention for their capabilities in multilingual natural language processing, while studies on risks associated with cross biases are limited to immediate context preferences. Cross-language…

Computation and Language · Computer Science 2025-08-07 Qianying Liu , Katrina Qiyao Wang , Fei Cheng , Sadao Kurohashi

Detection of hate speech has been formulated as a standalone application of NLP and different approaches have been adopted for identifying the target groups, obtaining raw data, defining the labeling process, choosing the detection…

Computation and Language · Computer Science 2023-09-07 Vitthal Bhandari

Recent developments in machine learning have shown that successful models do not rely only on huge amounts of data but the right kind of data. We show in this paper how this data-centric approach can be facilitated in a decentralized manner…

Computer Vision and Pattern Recognition · Computer Science 2022-10-31 M. R. Ahan , Robin Lehmann , Richard Blythman

Large Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective…

Computation and Language · Computer Science 2024-12-17 Tao Zhang , Ziqian Zeng , Yuxiang Xiao , Huiping Zhuang , Cen Chen , James Foulds , Shimei Pan

We investigate five English NLP benchmark datasets (on the superGLUE leaderboard) and two Swedish datasets for bias, along multiple axes. The datasets are the following: Boolean Question (Boolq), CommitmentBank (CB), Winograd Schema…

Computation and Language · Computer Science 2023-09-19 Tosin Adewumi , Isabella Södergren , Lama Alkhaled , Sana Sabah Sabry , Foteini Liwicki , Marcus Liwicki

We introduce new large labeled datasets on bias in 3 languages and show in experiments that bias exists in all 10 datasets of 5 languages evaluated, including benchmark datasets on the English GLUE/SuperGLUE leaderboards. The 3 new…

Computation and Language · Computer Science 2024-09-24 Irene Pagliai , Goya van Boven , Tosin Adewumi , Lama Alkhaled , Namrata Gurung , Isabella Södergren , Elisa Barney

Sophisticated language models such as OpenAI's GPT-3 can generate hateful text that targets marginalized groups. Given this capacity, we are interested in whether large language models can be used to identify hate speech and classify text…

Computation and Language · Computer Science 2022-03-25 Ke-Li Chiu , Annie Collins , Rohan Alexander

As large Pre-trained Language Models (PLMs) trained on large amounts of data in an unsupervised manner become more ubiquitous, identifying various types of bias in the text has come into sharp focus. Existing "Stereotype Detection" datasets…

Computation and Language · Computer Science 2022-03-29 Rajkumar Pujari , Erik Oveson , Priyanka Kulkarni , Elnaz Nouri

In this study, we investigate the capabilities and inherent biases of advanced large language models (LLMs) such as GPT-3.5 and GPT-4 in the context of debate evaluation. We discover that LLM's performance exceeds humans and surpasses the…

Computation and Language · Computer Science 2024-06-05 Xinyi Liu , Pinxin Liu , Hangfeng He

Instruction-tuned Large Language Models (LLMs) have recently showcased remarkable ability to generate fitting responses to natural language instructions. However, an open research question concerns the inherent biases of trained models and…

Computation and Language · Computer Science 2023-09-08 Patrick Haller , Ansar Aynetdinov , Alan Akbik

Recent studies have demonstrated how to assess the stereotypical bias in pre-trained English language models. In this work, we extend this branch of research in multiple different dimensions by systematically investigating (a) mono- and…

While deep learning models are making fast progress on the task of Natural Language Inference, recent studies have also shown that these models achieve high accuracy by exploiting several dataset biases, and without deep understanding of…

Computation and Language · Computer Science 2020-05-15 Xiang Zhou , Mohit Bansal

Large Language Models (LLMs) inherit societal biases from their training data, potentially leading to harmful or unfair outputs. While various techniques aim to mitigate these biases, their effects are often evaluated only along the…

Computation and Language · Computer Science 2025-11-25 Shireen Chand , Faith Baca , Emilio Ferrara

Prediction head is a crucial component of Transformer language models. Despite its direct impact on prediction, this component has often been overlooked in analyzing Transformers. In this study, we investigate the inner workings of the…

Computation and Language · Computer Science 2023-05-30 Goro Kobayashi , Tatsuki Kuribayashi , Sho Yokoi , Kentaro Inui

In a world increasingly reliant on artificial intelligence, it is more important than ever to consider the ethical implications of artificial intelligence on humanity. One key under-explored challenge is labeler bias, which can create…

Machine Learning · Computer Science 2024-10-25 Luke Haliburton , Sinksar Ghebremedhin , Robin Welsch , Albrecht Schmidt , Sven Mayer

Vision Language Models achieve impressive multi-modal performance but often inherit gender biases from their training data. This bias might be coming from both the vision and text modalities. In this work, we dissect the contributions of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Vivek Hruday Kavuri , Vysishtya Karanam , Venkata Jahnavi Venkamsetty , Kriti Madumadukala , Lakshmipathi Balaji Darur , Ponnurangam Kumaraguru

This study introduces a hypothesis-testing framework to assess whether large language models (LLMs) possess genuine reasoning abilities or primarily depend on token bias. We go beyond evaluating LLMs on accuracy; rather, we aim to…

Computation and Language · Computer Science 2024-10-07 Bowen Jiang , Yangxinyu Xie , Zhuoqun Hao , Xiaomeng Wang , Tanwi Mallick , Weijie J. Su , Camillo J. Taylor , Dan Roth