English

Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge

Computers and Society 2025-06-10 v2

Abstract

The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons" (Roose, 2024). In this position paper, we argue that the ML community would benefit from learning from and drawing on the social sciences when developing and using measurement instruments for evaluating GenAI systems. Specifically, our position is that evaluating GenAI systems is a social science measurement challenge. We present a four-level framework, grounded in measurement theory from the social sciences, for measuring concepts related to the capabilities, behaviors, and impacts of GenAI systems. This framework has two important implications: First, it can broaden the expertise involved in evaluating GenAI systems by enabling stakeholders with different perspectives to participate in conceptual debates. Second, it brings rigor to both conceptual and operational debates by offering a set of lenses for interrogating validity.

Keywords

Cite

@article{arxiv.2502.00561,
  title  = {Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge},
  author = {Hanna Wallach and Meera Desai and A. Feder Cooper and Angelina Wang and Chad Atalla and Solon Barocas and Su Lin Blodgett and Alexandra Chouldechova and Emily Corvi and P. Alex Dow and Jean Garcia-Gathright and Alexandra Olteanu and Nicholas Pangakis and Stefanie Reed and Emily Sheng and Dan Vann and Jennifer Wortman Vaughan and Matthew Vogel and Hannah Washington and Abigail Z. Jacobs},
  journal= {arXiv preprint arXiv:2502.00561},
  year   = {2025}
}

Comments

In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025

R2 v1 2026-06-28T21:29:10.324Z