English
Related papers

Related papers: InterFair: Debiasing with Natural Language Feedbac…

200 papers

Reward models (RMs) are essential in reinforcement learning from human feedback (RLHF) to align large language models (LLMs) with human values. However, RM training data is commonly recognized as low-quality, containing inductive biases…

Machine Learning · Computer Science 2026-05-20 Zhuo Li , Pengyu Cheng , Zhechao Yu , Feifei Tong , Anningzhe Gao , Tsung-Hui Chang , Xiang Wan , Erchao Zhao , Xiaoxi Jiang , Guanjun Jiang

Self-Correction based on feedback improves the output quality of Large Language Models (LLMs). Moreover, as Self-Correction functions like the slow and conscious System-2 thinking from cognitive psychology's perspective, it can potentially…

Computation and Language · Computer Science 2025-03-11 Panatchakorn Anantaprayoon , Masahiro Kaneko , Naoaki Okazaki

We investigate the prominent class of fair representation learning methods for bias mitigation. Using causal reasoning to define and formalise different sources of dataset bias, we reveal important implicit assumptions inherent to these…

Machine Learning · Computer Science 2025-02-11 Charles Jones , Fabio de Sousa Ribeiro , Mélanie Roschewitz , Daniel C. Castro , Ben Glocker

The question-answering (QA) capabilities of foundation models are highly sensitive to prompt variations, rendering their performance susceptible to superficial, non-meaning-altering changes. This vulnerability often stems from the model's…

Machine Learning · Computer Science 2024-06-07 Dyah Adila , Shuai Zhang , Boran Han , Yuyang Wang

Most works on gender bias focus on intrinsic bias -- removing traces of information about a protected group from the model's internal representation. However, these works are often disconnected from the impact of such debiasing on…

Computation and Language · Computer Science 2024-06-04 Bar Iluz , Yanai Elazar , Asaf Yehudai , Gabriel Stanovsky

What we discover and see online, and consequently our opinions and decisions, are becoming increasingly affected by automated machine learned predictions. Similarly, the predictive accuracy of learning machines heavily depends on the…

Information Retrieval · Computer Science 2020-01-15 Sami Khenissi , Olfa Nasraoui

Recent work has shown pre-trained language models capture social biases from the large amounts of text they are trained on. This has attracted attention to developing techniques that mitigate such biases. In this work, we perform an…

Computation and Language · Computer Science 2022-04-05 Nicholas Meade , Elinor Poole-Dayan , Siva Reddy

Good quality explanations strengthen the understanding of language models and data. Feature attribution methods, such as Integrated Gradient, are a type of post-hoc explainer that can provide token-level insights. However, explanations on…

Computation and Language · Computer Science 2026-04-21 Jonathan Kamp , Roos Bakker , Dominique Blok

A key assumption of most statistical machine learning methods is that they have access to independent samples from the distribution of data they encounter at test time. As such, these methods often perform poorly in the face of biased data,…

Machine Learning · Computer Science 2022-02-02 Xiaoting Shao , Karl Stelzner , Kristian Kersting

Extensive recent media focus has been directed towards the dark side of intelligent systems, how algorithms can influence society negatively. Often, transparency is proposed as a solution or step in the right direction. Unfortunately,…

Human-Computer Interaction · Computer Science 2018-12-11 Aaron Springer , Steve Whittaker

As machine learning algorithms are increasingly deployed for high-impact automated decision making, ethical and increasingly also legal standards demand that they treat all individuals fairly, without discrimination based on their age,…

Machine Learning · Computer Science 2021-05-03 Maarten Buyl , Tijl De Bie

Natural Language Processing (NLP) systems learn harmful societal biases that cause them to amplify inequality as they are deployed in more and more situations. To guide efforts at debiasing these systems, the NLP community relies on a…

Computation and Language · Computer Science 2021-06-09 Seraphina Goldfarb-Tarrant , Rebecca Marchant , Ricardo Muñoz Sanchez , Mugdha Pandya , Adam Lopez

Machine learning has become more important in real-life decision-making but people are concerned about the ethical problems it may bring when used improperly. Recent work brings the discussion of machine learning fairness into the causal…

Machine Learning · Statistics 2022-02-28 Haoyu Chen , Wenbin Lu , Rui Song , Pulak Ghosh

Model explanations such as saliency maps can improve user trust in AI by highlighting important features for a prediction. However, these become distorted and misleading when explaining predictions of images that are subject to systematic…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Wencan Zhang , Mariella Dimiccoli , Brian Y. Lim

The correct specification of reward models is a well-known challenge in reinforcement learning. Hand-crafted reward functions often lead to inefficient or suboptimal policies and may not be aligned with user values. Reinforcement learning…

Artificial Intelligence · Computer Science 2024-10-24 Muhan Lin , Shuyang Shi , Yue Guo , Behdad Chalaki , Vaishnav Tadiparthi , Ehsan Moradi Pari , Simon Stepputtis , Joseph Campbell , Katia Sycara

Personalization is pervasive in the online space as, when combined with learning, it leads to higher efficiency and revenue by allowing the most relevant content to be served to each user. However, recent studies suggest that such…

Computers and Society · Computer Science 2017-07-10 L. Elisa Celis , Nisheeth K. Vishnoi

Algorithmic decision-making systems are increasingly used throughout the public and private sectors to make important decisions or assist humans in making these decisions with real social consequences. While there has been substantial…

Human-Computer Interaction · Computer Science 2020-01-28 Ruotong Wang , F. Maxwell Harper , Haiyi Zhu

Prediction models can improve efficiency by automating decisions such as the approval of loan applications. However, they may inherit bias against protected groups from the data they are trained on. This paper adds counterfactual…

Machine Learning · Computer Science 2024-05-03 Nicholas Tenev

Training a model with access to human explanations can improve data efficiency and model performance on in- and out-of-domain data. Adding to these empirical findings, similarity with the process of human learning makes learning from…

Computation and Language · Computer Science 2022-04-20 Mareike Hartmann , Daniel Sonntag

Statistical natural language inference (NLI) models are susceptible to learning dataset bias: superficial cues that happen to associate with the label on a particular dataset, but are not useful in general, e.g., negation words indicate…

Computation and Language · Computer Science 2019-11-26 He He , Sheng Zha , Haohan Wang