English

Metric-DST: Mitigating Selection Bias Through Diversity-Guided Semi-Supervised Metric Learning

Machine Learning 2024-12-02 v2 Artificial Intelligence

Abstract

Selection bias poses a critical challenge for fairness in machine learning, as models trained on data that is less representative of the population might exhibit undesirable behavior for underrepresented profiles. Semi-supervised learning strategies like self-training can mitigate selection bias by incorporating unlabeled data into model training to gain further insight into the distribution of the population. However, conventional self-training seeks to include high-confidence data samples, which may reinforce existing model bias and compromise effectiveness. We propose Metric-DST, a diversity-guided self-training strategy that leverages metric learning and its implicit embedding space to counter confidence-based bias through the inclusion of more diverse samples. Metric-DST learned more robust models in the presence of selection bias for generated and real-world datasets with induced bias, as well as a molecular biology prediction task with intrinsic bias. The Metric-DST learning strategy offers a flexible and widely applicable solution to mitigate selection bias and enhance fairness of machine learning models.

Keywords

Cite

@article{arxiv.2411.18442,
  title  = {Metric-DST: Mitigating Selection Bias Through Diversity-Guided Semi-Supervised Metric Learning},
  author = {Yasin I. Tepeli and Mathijs de Wolf and Joana P. Gonçalves},
  journal= {arXiv preprint arXiv:2411.18442},
  year   = {2024}
}

Comments

18 pages main manuscript (4 main figures), 7 pages of supplementary

R2 v1 2026-06-28T20:14:44.307Z