理解算法公平性细分评估interpretation中的挑战
机器学习
2025-10-27 v2 计算机与社会
机器学习
摘要
跨子群的细分评估对于评估机器学习模型的公平性至关重要,但其不加批判性的使用可能误导实践者。我们表明,在数据代表相关人口但反映现实差异的情况下,子群间的等性能是一个不可靠的公平性衡量标准。此外,当数据因选择偏差而不具代表性时,无论是细分评估还是基于条件独立性测试的替代方法,都可能在不明确假设偏差机制的情况下失效。我们使用因果图模型来描述在不同数据生成过程下公平性属性和指标稳定性。我们的框架建议在控制混杂和分布迁移(包括条件独立测试和加权性能估计)后,补充细分评估中的显式因果假设和分析。这些发现对实践者在给定细分评估普遍性的情况下设计和解释模型评估具有广泛影响。
引用
@article{arxiv.2506.04193,
title = {Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness},
author = {Stephen R. Pfohl and Natalie Harris and Chirag Nagpal and David Madras and Vishwali Mhasawade and Olawale Salaudeen and Awa Dieng and Shannon Sequeira and Santiago Arciniegas and Lillian Sung and Nnamdi Ezeanochie and Heather Cole-Lewis and Katherine Heller and Sanmi Koyejo and Alexander D'Amour},
journal= {arXiv preprint arXiv:2506.04193},
year = {2025}
}
备注
To be published at the 2025 Conference on Neural Information Processing Systems (NeurIPS 2025)