English

Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation

Computation and Language 2026-05-01 v1 Artificial Intelligence

Abstract

We adapted the Reliable Change Index (RCI; Jacobson and Truax, 1991) from clinical psychology to item-level LLM version comparison on 2,000 MMLU-Pro items (K=10 samples at T=0.7). Two within-family pairs were tested: Llama 3 to 3.1 (+1.6 points) and Qwen 2.5 to 3 (+2.8 points). On the full benchmark, most items showed no reliable change (79% and 72%). However, over half the items were floor/ceiling. Among analysable items, change was bidirectional with large effect sizes: 34% improved and 28% deteriorated for Llama; 47% improved and 39% deteriorated for Qwen (median |delta p| = 0.50 and 0.90). Churn was asymmetric by difficulty: low-accuracy items improved, high-accuracy items deteriorated. Domain-level decomposition revealed family-specific reversals: Llama lost physics while Qwen lost law. Greedy single-shot evaluation missed 42% of reliably changed items and falsely flagged 25% of unchanged items. The aggregate accuracy gain is the net residual of opposing item-level movements. We recommend reporting churn rate alongside aggregate accuracy.

Keywords

Cite

@article{arxiv.2604.27405,
  title  = {Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation},
  author = {Jon-Paul Cacioli},
  journal= {arXiv preprint arXiv:2604.27405},
  year   = {2026}
}

Comments

7 pages, 4 figures, 2 tables. Pre-registered study. Code and data available

R2 v1 2026-07-01T12:42:52.223Z