English

From Human-Level AI Tales to AI Leveling Human Scales

Machine Learning 2026-04-08 v2

Abstract

Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose a framework that calibrates items against the 'world population' and report performance on a common, human-anchored scale. Concretely, we build on a set of multi-level scales for different capabilities where each level should represent a probability of success of the whole world population on a logarithmic scale with a base BB. We calibrate each scale for each capability (reasoning, comprehension, knowledge, volume, etc.) by compiling publicly released human test data spanning education and reasoning benchmarks (PISA, TIMSS, ICAR, UKBioBank, and ReliabilityBench). The base BB is estimated by extrapolating between samples with two demographic profiles using LLMs, with the hypothesis that they condense rich information about human populations. We evaluate the quality of different mappings using group slicing and post-stratification. The new techniques allow for the recalibration and standardization of scales relative to the whole-world population.

Keywords

Cite

@article{arxiv.2602.18911,
  title  = {From Human-Level AI Tales to AI Leveling Human Scales},
  author = {Peter Romero and Fernando Martínez-Plumed and Zachary R. Tidler and Matthieu Téhénan and Sipeng Chen and Álvaro David Gómez Antón and Luning Sun and Manuel Cebrian and Lexin Zhou and Yael Moros Daval and Daniel Romero-Alvarado and Félix Martí Pérez and Kevin Wei and José Hernández-Orallo},
  journal= {arXiv preprint arXiv:2602.18911},
  year   = {2026}
}

Comments

23 pages, 10 figures. submitted to ICML 2026