English
Related papers

Related papers: Reliability Gaps Between Groups in COMPAS Dataset

200 papers

Despite the constant development of new bias mitigation methods for machine learning, no method consistently succeeds, and a fundamental question remains unanswered: when and why do bias mitigation techniques fail? In this paper, we…

Machine Learning · Computer Science 2025-07-15 Anissa Alloula , Charles Jones , Ben Glocker , Bartłomiej W. Papież

Recently, it was shown that most popular IR measures are not interval-scaled, implying that decades of experimental IR research used potentially improper methods, which may have produced questionable results. However, it was unclear if and…

Information Retrieval · Computer Science 2021-01-08 Marco Ferrante , Nicola Ferro , Norbert Fuhr

Causal discovery can be a powerful tool for investigating causality when a system can be observed but is inaccessible to experiments in practice. Despite this, it is rarely used in any scientific or medical fields. One of the major hurdles…

Machine Learning · Statistics 2019-10-07 Erich Kummerfeld , Alexander Rix

Non-reference speech quality models are important for a growing number of applications. The VoiceMOS 2022 challenge provided a dataset of synthetic voice conversion and text-to-speech samples with subjective labels. This study looks at the…

Sound · Computer Science 2022-09-15 Michael Chinen , Jan Skoglund , Chandan K A Reddy , Alessandro Ragano , Andrew Hines

Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on medical imaging…

Mathematical models of real life phenomena are highly nonlinear involving multiple parameters and often exhibiting complex dynamics. Experimental data sets are typically small and noisy, rendering estimation of parameters from such data…

Chaotic Dynamics · Physics 2017-05-11 Abhirup Ghosh , Samit Bhattacharyya , Somdatta Sinha , Amit Apte

Disaggregated evaluation across subgroups is critical for assessing the fairness of machine learning models, but its uncritical use can mislead practitioners. We show that equal performance across subgroups is an unreliable measure of…

Randomized experiments have become a cornerstone of evidence-based decision-making in contexts ranging from online platforms to public health. However, in experimental settings with network interference, a unit's treatment can influence…

Machine Learning · Computer Science 2025-10-22 Sadegh Shirani , Yuwei Luo , William Overman , Ruoxuan Xiong , Mohsen Bayati

We provide the asymptotic distribution of the major indexes used in the statistical literature to quantify disparate treatment in machine learning. We aim at promoting the use of confidence intervals when testing the so-called group…

Machine Learning · Statistics 2018-07-18 Philippe Besse , Eustasio del Barrio , Paula Gordaliza , Jean-Michel Loubes

Although the widespread use of AI systems in today's world is growing, many current AI systems are found vulnerable due to hidden bias and missing information, especially in the most commonly used forecasting system. In this work, we…

Machine Learning · Computer Science 2024-07-30 Zhixuan Chu , Hui Ding , Guang Zeng , Shiyu Wang , Yiming Li

Dataset replication is a useful tool for assessing whether improvements in test accuracy on a specific benchmark correspond to improvements in models' ability to generalize reliably. In this work, we present unintuitive yet significant ways…

Machine learning models only provide probabilistic guarantees on the expected loss of random samples from the distribution represented by their training data. As a result, a model with high accuracy, may or may not be reliable for…

Databases · Computer Science 2024-04-12 Nima Shahbazi , Abolfazl Asudeh

We provide a comprehensive analysis of the differences between two important standards for randomized benchmarking (RB): the Clifford-group RB protocol proposed originally in Emerson et al (2005) and Dankert et al (2006), and a variant of…

Quantum Physics · Physics 2019-03-27 Kristine Boone , Arnaud Carignan-Dugas , Joel J. Wallman , Joseph Emerson

The acoustic variability of noisy and reverberant speech mixtures is influenced by multiple factors, such as the spectro-temporal characteristics of the target speaker and the interfering noise, the signal-to-noise ratio (SNR) and the room…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-09 Philippe Gonzalez , Tommy Sonne Alstrøm , Tobias May

As AI becomes prevalent in high-risk domains and decision-making, it is essential to test for potential harms and biases. This urgency is reflected by the global emergence of AI regulations that emphasise fairness and adequate testing, with…

Machine Learning · Computer Science 2025-07-25 Varsha Ramineni , Hossein A. Rahmani , Emine Yilmaz , David Barber

Language technologies have a racial bias, committing greater errors for Black users than for white users. However, little work has evaluated what effect these disparate error rates have on users themselves. The present study aims to…

Human-Computer Interaction · Computer Science 2023-02-27 Kimi Wenzel , Nitya Devireddy , Cam Davidson , Geoff Kaufman

In this paper, we discuss causal inference on the efficacy of a treatment or medication on a time-to-event outcome with competing risks. Although the treatment group can be randomized, there can be confoundings between the compliance and…

Methodology · Statistics 2016-12-06 Cheng Zheng , Ran Dai , Parameswaran Hari , Mei-Jie Zhang

This study examines opinion instability among individuals from different ethnic groups (White, Latino, and Asian Americans) by analyzing measurement errors in survey measures. Using a multi-wave panel dataset of college students and…

Applications · Statistics 2024-06-10 Bang Quan Zheng

With the proliferation of algorithmic decision-making, increased scrutiny has been placed on these systems. This paper explores the relationship between the quality of the training data and the overall fairness of the models trained with…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Aki Barry , Lei Han , Gianluca Demartini

Individual and social biases undermine the effectiveness of human advisers by inducing judgment errors which can disadvantage protected groups. In this paper, we study the influence these biases can have in the pervasive problem of fake…

Human-Computer Interaction · Computer Science 2024-03-15 Axel Abels , Elias Fernandez Domingos , Ann Nowé , Tom Lenaerts