检索中的多群组比例代表性
人工智能
2024-11-04 v2 信息检索
信息论
机器学习
math.IT
机器学习
摘要
图像搜索与检索任务可能延续有害刻板印象,抹去文化身份,并放大社会差异。当前用于缓解此类代表性伤害的方法通过平衡由少量(常为二元)属性定义的族群中检索物品的数量来实现,但大多数现有方法忽略了由群体属性组合(如性别、种族和族裔)所决定的交叉族群。我们提出一种新型指标——多群组比例代表性(MPR),用于衡量跨交叉族群的代表性。我们发展了实用的方法来估计 MPR,提供了理论保证,并提出了确保检索中 MPR 的优化算法。我们展示,优化等比例代表性指标的现有方法可能无法促进 MPR。关键的是,我们的工作表明,优化 MPR 可在通常以最小检索精度损失为代价的情况下,实现对由丰富函数类指定的多个交叉族群的更比例化代表性。
引用
@article{arxiv.2407.08571,
title = {Multi-Group Proportional Representation in Retrieval},
author = {Alex Oesterling and Claudio Mayrink Verdun and Carol Xuan Long and Alexander Glynn and Lucas Monteiro Paes and Sajani Vithana and Martina Cardone and Flavio P. Calmon},
journal= {arXiv preprint arXiv:2407.08571},
year = {2024}
}
备注
48 pages, 33 figures. Accepted as poster at NeurIPS 2024. Code can be found at https://github.com/alex-oesterling/multigroup-proportional-representation