面向公平性与隐私:一种用于非二元受保护属性的数据预处理优化框架
摘要
AI 不公平结果的原因常常根植于偏差数据集。因此,本工作提出了一种框架,用于通过去偏处理包含(非)二元受保护属性的数据集来解决公平性问题。该框架提出了一个组合优化问题,可使用诸如遗传算法等启发式方法来求解所述公平性目标。该框架通过寻找最小化某种歧视度量的数据子集来实现这一点。 Depending on a user-defined setting, the framework enables different use cases, such as data removal, the addition of synthetic data, or exclusive use of synthetic data. The exclusive use of synthetic data in particular enhances the framework's ability to preserve privacy while optimizing for fairness. In a comprehensive evaluation, we demonstrate that under our framework, genetic algorithms can effectively yield fairer datasets compared to the original data. In contrast to prior work, the framework exhibits a high degree of flexibility as it is metric- and task-agnostic, can be applied to both binary or non-binary protected attributes, and demonstrates efficient runtime.
引用
@article{arxiv.2410.00836,
title = {Towards Fairness and Privacy: A Novel Data Pre-processing Optimization Framework for Non-binary Protected Attributes},
author = {Manh Khoi Duong and Stefan Conrad},
journal= {arXiv preprint arXiv:2410.00836},
year = {2024}
}
备注
The Version of Record of this contribution is published in Data Science and Machine Learning, volume 1943, CCIS (Springer Singapore) 2023. It is available online at https://doi.org/10.1007/978-981-99-8696-5