中文

模型评估中严谨透明人类基线的建议与报告清单

人工智能 2025-11-04 v1 计算机与社会 人机交互

摘要

在这篇立场论文中,我们认为基础模型评估中的人类基线必须更加严谨和透明,以实现人类与AI性能的有意义比较,并为此提供建议和一份报告清单。人类性能基线对于机器学习社区、下游用户和政策制定者解读AI评估至关重要。模型常被声称达到“超人”性能,但现有的基线方法既不够严谨,也未被充分记录,无法稳健地测量和评估性能差异。基于对测量理论和AI评估文献的元回顾,我们推导出一个框架,为设计、执行和报告人类基线提供建议。我们将建议综合成一份清单,并用于系统性地审查基础模型评估中的115个人类基线(研究),从而识别出现有基线方法的缺陷;我们的清单也可以帮助研究人员进行人类基线并报告结果。我们希望我们的工作能够推进更严谨的AI评估实践,更好地服务于研究社区和政策制定者。数据可在 https://github.com/kevinlwei/human-baselines 获取。

关键词

引用

@article{arxiv.2506.13776,
  title  = {Recommendations and Reporting Checklist for Rigorous & Transparent Human Baselines in Model Evaluations},
  author = {Kevin L. Wei and Patricia Paskov and Sunishchal Dev and Michael J. Byun and Anka Reuel and Xavier Roberts-Gaal and Rachel Calcott and Evie Coxon and Chinmay Deshpande},
  journal= {arXiv preprint arXiv:2506.13776},
  year   = {2025}
}

备注

A version of this paper has been accepted to ICML 2025 as a position paper (spotlight), with the title: "Position: Human Baselines in Model Evaluations Need Rigor and Transparency (With Recommendations & Reporting Checklist)."