面向机器学习模型的多利益相关者评估:关于职位匹配系统中指标偏好的众包研究
计算机与社会
2025-03-11 v1 人工智能
摘要
尽管机器学习(ML)技术影响着 diverse 利益相关者,但并不存在一种放之四海而皆准的指标来评估其输出的质量,包括性能与公平性。在不征求利益相关者意见的情况下使用预设指标是有问题的,因为这会导致在 ML 流程中对利益相关者的不公平忽视。在本研究中,为了建立将不同利益相关者意见纳入 ML 指标选择的实用方法,我们利用众包技术调查了参与者对不同指标的偏好。我们要求 837 名参与者在假设的职位匹配系统中,从两个假设的 ML 模型中选择更好的模型,共进行二十次,并计算了他们在七个指标上的效用值。为了详细检查参与者的反馈,我们根据其效用值将他们分为五个聚类,并分析了每个聚类的倾向,包括他们对指标的偏好和共同属性。基于这些结果,我们讨论了在为多个利益相关者选择适当指标和评估 ML 模型时应考虑的要点。
引用
@article{arxiv.2503.05796,
title = {Towards Multi-Stakeholder Evaluation of ML Models: A Crowdsourcing Study on Metric Preferences in Job-matching System},
author = {Takuya Yokota and Yuri Nakao},
journal= {arXiv preprint arXiv:2503.05796},
year = {2025}
}
备注
This version of the contribution has been accepted for publication, after peer review (when applicable) but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. Use of this Accepted Version is subject to the publisher's Accepted Manuscript terms of use https://www.springernature.com/gp/open-research/policies/accepted-manuscript-terms