English
Related papers

Related papers: Gender Bias in Explainability: Investigating Perfo…

200 papers

Post hoc explanations have emerged as a way to improve user trust in machine learning models by providing insight into model decision-making. However, explanations tend to be evaluated based on their alignment with prior knowledge while the…

Human-Computer Interaction · Computer Science 2023-12-13 Tessa Han , Yasha Ektefaie , Maha Farhat , Marinka Zitnik , Himabindu Lakkaraju

Gender, race and social biases have recently been detected as evident examples of unfairness in applications of Natural Language Processing. A key path towards fairness is to understand, analyse and interpret our data and algorithms. Recent…

Computation and Language · Computer Science 2021-05-06 Christine Basta , Marta R. Costa-jussà

Algorithmic systems such as search engines and information retrieval platforms significantly influence academic visibility and the dissemination of knowledge. Despite assumptions of neutrality, these systems can reproduce or reinforce…

Information Retrieval · Computer Science 2025-08-11 Stefanie Urchs , Veronika Thurner , Matthias Aßenmacher , Ludwig Bothmann , Christian Heumann , Stephanie Thiemichen

Moral alignment has emerged as a widely adopted approach for regulating the behavior of pretrained language models (PLMs), typically through fine-tuning on curated datasets. Gender stereotype mitigation is a representational task within the…

Computation and Language · Computer Science 2025-11-21 Guangliang Liu , Bocheng Chen , Han Zi , Xitong Zhang , Kristen Marie Johnson

AI systems have been known to amplify biases in real-world data. Explanations may help human-AI teams address these biases for fairer decision-making. Typically, explanations focus on salient input features. If a model is biased against…

Artificial Intelligence · Computer Science 2024-04-10 Navita Goyal , Connor Baumler , Tin Nguyen , Hal Daumé

Post-hoc explanation techniques refer to a posteriori methods that can be used to explain how black-box machine learning models produce their outcomes. Among post-hoc explanation techniques, counterfactual explanations are becoming one of…

Machine Learning · Computer Science 2020-09-07 Ulrich Aïvodji , Alexandre Bolot , Sébastien Gambs

Large Language Models (LLMs) have revolutionized natural language processing, yet concerns persist regarding their tendency to reflect or amplify social biases. This study introduces a novel evaluation framework to uncover gender biases in…

Computation and Language · Computer Science 2026-03-10 Evan Chen , Run-Jun Zhan , Yan-Bai Lin , Hung-Hsuan Chen

Large Language Models (LLMs) can generate biased responses. Yet previous direct probing techniques contain either gender mentions or predefined gender stereotypes, which are challenging to comprehensively collect. Hence, we propose an…

Computation and Language · Computer Science 2024-02-20 Xiangjue Dong , Yibo Wang , Philip S. Yu , James Caverlee

Machine learning and deep learning models are pivotal in educational contexts, particularly in predicting student success. Despite their widespread application, a significant gap persists in comprehending the factors influencing these…

Machine Learning · Computer Science 2024-05-24 Priscylla Silva , Claudio T. Silva , Luis Gustavo Nonato

This study investigates factors influencing Automatic Speech Recognition (ASR) systems' fairness and performance across genders, beyond the conventional examination of demographics. Using the LibriSpeech dataset and the Whisper small model,…

Computation and Language · Computer Science 2025-02-26 Hend ElGhazaly , Bahman Mirheidari , Nafise Sadat Moosavi , Heidi Christensen

We investigate the prominent class of fair representation learning methods for bias mitigation. Using causal reasoning to define and formalise different sources of dataset bias, we reveal important implicit assumptions inherent to these…

Machine Learning · Computer Science 2025-02-11 Charles Jones , Fabio de Sousa Ribeiro , Mélanie Roschewitz , Daniel C. Castro , Ben Glocker

A multitude of explainability methods and associated fidelity performance metrics have been proposed to help better understand how modern AI systems make decisions. However, much of the current work has remained theoretical -- without much…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Julien Colin , Thomas Fel , Remi Cadene , Thomas Serre

Existing and planned legislation stipulates various obligations to provide information about machine learning algorithms and their functioning, often interpreted as obligations to "explain". Many researchers suggest using post-hoc…

Machine Learning · Computer Science 2022-05-12 Sebastian Bordt , Michèle Finck , Eric Raidl , Ulrike von Luxburg

The proliferation of personalized recommendation technologies has raised concerns about discrepancies in their recommendation performance across different genders, age groups, and racial or ethnic populations. This varying degree of…

Information Retrieval · Computer Science 2020-02-19 Masoud Mansoury , Himan Abdollahpouri , Jessie Smith , Arman Dehpanah , Mykola Pechenizkiy , Bamshad Mobasher

The rise of machine learning (ML) is accompanied by several high-profile cases that have stressed the need for fairness, accountability, explainability and trust in ML systems. The existing literature has largely focused on fully automated…

Computers and Society · Computer Science 2023-06-14 Bhavya Ghai

Problem. Educational disparities in Mathematics performance are a persistent challenge. This study aims to unravel the complex factors contributing to these disparities among students internationally, with a focus on the interpretability of…

Computers and Society · Computer Science 2025-02-28 Ismael Gomez-Talal , Luis Bote-Curiel , Jose Luis Rojo-Alvarez

We investigate whether three types of post hoc model explanations--feature attribution, concept activation, and training point ranking--are effective for detecting a model's reliance on spurious signals in the training data. Specifically,…

Machine Learning · Computer Science 2022-12-12 Julius Adebayo , Michael Muelly , Hal Abelson , Been Kim

Trustworthy machine learning in healthcare requires strong predictive performance, fairness, and explanations. While it is known that improving fairness can affect predictive performance, little is known about how fairness improvements…

Machine Learning · Computer Science 2025-12-03 Joshua Wolff Anderson , Shyam Visweswaran

Post-hoc interpretability methods play a critical role in explainable artificial intelligence (XAI), as they pinpoint portions of data that a trained deep learning model deemed important to make a decision. However, different post-hoc…

Machine Learning · Computer Science 2024-07-30 Jiawen Wei , Hugues Turbé , Gianmarco Mengaldo

Deploying an algorithmically informed policy is a significant intervention in the structure of society. As is increasingly acknowledged, predictive algorithms have performative effects: using them can shift the distribution of social…

Computers and Society · Computer Science 2025-05-01 Sebastian Zezulka , Konstantin Genin