English
Related papers

Related papers: ProPublica's COMPAS Data Revisited

200 papers

Observational cohort studies with oversampled exposed subjects are typically implemented to understand the causal effect of a rare exposure. Because the distribution of exposed subjects in the sample differs from the source population,…

Methodology · Statistics 2019-02-14 Sherri Rose

Machine learning algorithms permeate the day-to-day aspects of our lives and therefore studying the fairness of these algorithms before implementation is crucial. One way in which bias can manifest in a dataset is through missing values.…

Machine Learning · Statistics 2026-02-23 Aeysha Bhatti , Trudie Sandrock , Johane Nienkemper-Swanepoel

Recently, there has been a rising awareness that when machine learning (ML) algorithms are used to automate choices, they may treat/affect individuals unfairly, with legal, ethical, or economic consequences. Recommender systems are…

Information Retrieval · Computer Science 2022-04-19 Mohammadmehdi Naghiaei , Hossein A. Rahmani , Yashar Deldjoo

Despite an increasing reliance on fully-automated algorithmic decision-making in our day-to-day lives, human beings still make highly consequential decisions. As frequently seen in business, healthcare, and public policy, recommendations…

Computers and Society · Computer Science 2021-12-14 Kosuke Imai , Zhichao Jiang , James Greiner , Ryan Halen , Sooahn Shin

Sentiment signals derived from sparse news are commonly used in financial analysis and technology monitoring, yet transforming raw article-level observations into reliable temporal series remains a largely unsolved engineering problem.…

Machine Learning · Computer Science 2026-03-26 Stefania Stan , Marzio Lunghi , Vito Vargetto , Claudio Ricci , Rolands Repetto , Brayden Leo , Shao-Hong Gan

As calls for fair and unbiased algorithmic systems increase, so too does the number of individuals working on algorithmic fairness in industry. However, these practitioners often do not have access to the demographic data they feel they…

Computers and Society · Computer Science 2021-01-26 McKane Andrus , Elena Spitzer , Jeffrey Brown , Alice Xiang

Synthetic data is increasingly used to support research without exposing sensitive user content. Social media data is one of the types of datasets that would hugely benefit from representative synthetic equivalents that can be used to…

Cryptography and Security · Computer Science 2026-03-06 Henry Tari , Adriana Iamnitchi

Deep learning-based facial recognition systems have experienced increased media attention due to exhibiting unfair behavior. Large enterprises, such as IBM, shut down their facial recognition and age prediction systems as a consequence. Age…

Computer Vision and Pattern Recognition · Computer Science 2021-11-17 Yushi Cao , David Berend , Palina Tolmach , Guy Amit , Moshe Levy , Yang Liu , Asaf Shabtai , Yuval Elovici

Adequate sampling space coverage is the keystone to effectively train trustworthy Machine Learning models. Unfortunately, real data do carry several inherent risks due to the many potential biases they exhibit when gathered without a proper…

Machine Learning · Computer Science 2025-03-27 Antonio Maratea , Rita Perna

Data-driven predictive algorithms are widely used to automate and guide high-stake decision making such as bail and parole recommendation, medical resource distribution, and mortgage allocation. Nevertheless, harmful outcomes biased against…

Computers and Society · Computer Science 2022-06-03 Atoosa Kasirzadeh

With promising empirical performance across a wide range of applications, synthetic data augmentation appears a viable solution to data scarcity and the demands of increasingly data-intensive models. Its effectiveness lies in expanding the…

Machine Learning · Computer Science 2026-02-02 Zixuan Wu , So Won Jeong , Yating Liu , Yeo Jin Jung , Claire Donnat

Causal inference in longitudinal biomedical data remains a central challenge, especially in psychiatry, where symptom heterogeneity and latent confounding frequently undermine classical estimators. Most existing methods for treatment effect…

Machine Learning · Computer Science 2025-07-28 Eric V. Strobl

In the field of data mining, how to deal with high-dimensional data is an inevitable problem. Unsupervised feature selection has attracted more and more attention because it does not rely on labels. The performance of spectral-based…

Machine Learning · Computer Science 2021-01-01 Zhengxin Li , Feiping Nie , Jintang Bian , Xuelong Li

Although the fairness community has recognized the importance of data, researchers in the area primarily rely on UCI Adult when it comes to tabular data. Derived from a 1994 US Census survey, this dataset has appeared in hundreds of…

Machine Learning · Computer Science 2022-01-11 Frances Ding , Moritz Hardt , John Miller , Ludwig Schmidt

The discovery of discriminatory bias in human or automated decision making is a task of increasing importance and difficulty, exacerbated by the pervasive use of machine learning and data mining. Currently, discrimination discovery largely…

Computers and Society · Computer Science 2019-11-05 Bilal Qureshi , Faisal Kamiran , Asim Karim , Salvatore Ruggieri , Dino Pedreschi

The use of algorithmic (learning-based) decision making in scenarios that affect human lives has motivated a number of recent studies to investigate such decision making systems for potential unfairness, such as discrimination against…

Machine Learning · Computer Science 2021-05-11 Junaid Ali , Muhammad Bilal Zafar , Adish Singla , Krishna P. Gummadi

We study a data analyst's problem of acquiring data from self-interested individuals to obtain an accurate estimation of some statistic of a population, subject to an expected budget constraint. Each data holder incurs a cost, which is…

Computer Science and Game Theory · Computer Science 2019-05-15 Yiling Chen , Shuran Zheng

Big data and algorithmic risk prediction tools promise to improve criminal justice systems by reducing human biases and inconsistencies in decision making. Yet different, equally-justifiable choices when developing, testing, and deploying…

Computers and Society · Computer Science 2022-09-23 Travis Greene , Galit Shmueli , Jan Fell , Ching-Fu Lin , Han-Wei Liu

Data-driven algorithms play a large role in decision making across a variety of industries. Increasingly, these algorithms are being used to make decisions that have significant ramifications for people's social and economic well-being,…

Machine Learning · Computer Science 2018-09-26 J. Henry Hinnefeld , Peter Cooman , Nat Mammo , Rupert Deese

Propensity scores are commonly used to estimate treatment effects from observational data. We argue that the probabilistic output of a learned propensity score model should be calibrated -- i.e., a predictive treatment probability of 90%…

Methodology · Statistics 2024-06-06 Shachi Deshpande , Volodymyr Kuleshov