English

Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scores

Artificial Intelligence 2025-09-30 v1

Abstract

New large language models (LLMs) are being released every day. Some perform significantly better or worse than expected given their parameter count. Therefore, there is a need for a method to independently evaluate models. The current best way to evaluate a model is to measure its Elo score by comparing it to other models in a series of contests - an expensive operation since humans are ideally required to compare LLM outputs. We observe that when an LLM is asked to judge such contests, the consistency with which it selects a model as the best in a matchup produces a metric that is 91% correlated with its own human-produced Elo score. This provides a simple proxy for Elo scores that can be computed cheaply, without any human data or prior knowledge.

Keywords

Cite

@article{arxiv.2509.23510,
  title  = {Model Consistency as a Cheap yet Predictive Proxy for LLM Elo Scores},
  author = {Ashwin Ramaswamy and Nestor Demeure and Ermal Rrapaj},
  journal= {arXiv preprint arXiv:2509.23510},
  year   = {2025}
}
R2 v1 2026-07-01T06:01:35.178Z