English
Related papers

Related papers: Elo Uncovered: Robustness and Best Practices in La…

200 papers

Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Open-ended evaluation systems, where candidate models are…

Computer Science and Game Theory · Computer Science 2025-05-09 Siqi Liu , Ian Gemp , Luke Marris , Georgios Piliouras , Nicolas Heess , Marc Lanctot

Elo rating, widely used for skill assessment across diverse domains ranging from competitive games to large language models, is often understood as an incremental update algorithm for estimating a stationary Bradley-Terry (BT) model.…

Machine Learning · Computer Science 2025-02-18 Shange Tang , Yuanhao Wang , Chi Jin

Deciding which large language model (LLM) to use is a complex challenge. Pairwise ranking has emerged as a new method for evaluating human preferences for LLMs. This approach entails humans evaluating pairs of model outputs based on a…

Computation and Language · Computer Science 2025-02-18 Roland Daynauth , Christopher Clarke , Krisztian Flautner , Lingjia Tang , Jason Mars

We introduce LLM CHESS, an evaluation framework designed to probe the generalization of reasoning and instruction-following abilities in large language models (LLMs) through extended agentic interaction in the domain of chess. We rank over…

Artificial Intelligence · Computer Science 2025-12-02 Sai Kolasani , Maxim Saplin , Nicholas Crispino , Kyle Montgomery , Jared Quincy Davis , Matei Zaharia , Chi Wang , Chenguang Wang

As large language models (LLMs) continue to advance, accurately and comprehensively evaluating their performance becomes increasingly challenging. Ranking the relative performance of LLMs based on Elo ratings, according to human judgment,…

Computation and Language · Computer Science 2023-11-14 Minghao Wu , Alham Fikri Aji

New large language models (LLMs) are being released every day. Some perform significantly better or worse than expected given their parameter count. Therefore, there is a need for a method to independently evaluate models. The current best…

Artificial Intelligence · Computer Science 2025-09-30 Ashwin Ramaswamy , Nestor Demeure , Ermal Rrapaj

Rating systems play a crucial role in evaluating player skill across competitive environments. The Elo rating system, originally designed for deterministic and information-complete games such as chess, has been widely adopted and modified…

Computer Science and Game Theory · Computer Science 2025-12-23 Avirup Chakraborty , Shirsa Maitra , Tathagata Banerjee , Diganta Mukherjee , Tridib Mukherjee

Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical challenge due to data contamination, opaque operation, and subjective preferences. To address…

Artificial Intelligence · Computer Science 2026-04-15 Qianhong Guo , Wei Xie , Xiaofang Cai , Enze Wang , Shuoyoucheng Ma , Xiaobing Sun , Tian Xia , Kai Chen , Xiaofeng Wang , Baosheng Wang

As Large Language Models (LLMs) achieve breakthroughs in complex reasoning, Codeforces-based Elo ratings have emerged as a prominent metric for evaluating competitive programming capabilities. However, these ratings are often reported…

Software Engineering · Computer Science 2026-02-06 Shenyu Zheng , Ximing Dong , Xiaoshuang Liu , Gustavo Oliva , Chong Chun Yong , Dayi Lin , Boyuan Chen , Shaowei Wang , Ahmed E. Hassan

The Elo rating system is widely adopted to evaluate the skills of (chess) game and sports players. Recently it has been also integrated into machine learning algorithms in evaluating the performance of computerised AI agents. However, an…

Machine Learning · Computer Science 2022-01-21 Xue Yan , Yali Du , Binxin Ru , Jun Wang , Haifeng Zhang , Xu Chen

The TextClass Benchmark project is an ongoing, continuous benchmarking process that aims to provide a comprehensive, fair, and dynamic evaluation of LLMs and transformers for text classification tasks. This evaluation spans various domains…

Computation and Language · Computer Science 2024-12-10 Bastián González-Bustamante

The evaluation of open-ended responses in serious games presents a unique challenge, as correctness is often subjective. Large Language Models (LLMs) are increasingly being explored as evaluators in such contexts, yet their accuracy and…

Computation and Language · Computer Science 2025-04-18 Andrés Isaza-Giraldo , Paulo Bala , Lucas Pereira

The Elo rating system, which was originally proposed by Arpad Elo for chess, has become one of the most important rating systems in sports, economics and gaming nowadays. Its original formulation is based on two-player zero-sum games, but…

Optimization and Control · Mathematics 2022-04-12 Düring Bertram , Fischer Michael , Wolfram Marie-Therese

In this work, we explore the Large Language Model (LLM) agent reviewer dynamics in an Elo-ranked review system using real-world conference paper submissions. Multiple LLM agent reviewers with different personas are engage in multi round…

Computation and Language · Computer Science 2026-01-14 Hsiang-Wei Huang , Junbin Lu , Kuang-Ming Chen , Jenq-Neng Hwang

Large language models (LLMs) have been extensively used as the backbones for general-purpose agents, and some economics literature suggest that LLMs are capable of playing various types of economics games. Following these works, to overcome…

Computer Science and Game Theory · Computer Science 2024-01-04 Shangmin Guo , Haoran Bu , Haochuan Wang , Yi Ren , Dianbo Sui , Yuming Shang , Siting Lu

It was recently observed that Elo ratings fail at preserving transitive relations among strategies and therefore cannot correctly extract the transitive component of a game. We provide a characterization of transitive games as a weak…

Computer Science and Game Theory · Computer Science 2024-03-07 Nelson Vadori , Rahul Savani

While Large Language Models (LLMs) are fundamentally next-token prediction systems, their practical applications extend far beyond this basic function. From natural language processing and text generation to conversational assistants and…

Computation and Language · Computer Science 2025-03-10 Vishakha Agrawal , Archie Chaudhury , Shreya Agrawal

Ideal or real - that is the question.In this work, we explore whether principles from game theory can be effectively applied to the evaluation of large language models (LLMs). This inquiry is motivated by the growing inadequacy of…

Computation and Language · Computer Science 2026-04-07 Gao Yang , Yuhang Liu , Siyu Miao , Xinyue Liang , Zhengyang Liu , Heyan Huang

The Elo algorithm, renowned for its simplicity, is widely used for rating in sports tournaments and other applications. However, despite its widespread use, a detailed understanding of the convergence characteristics of the Elo algorithm is…

Machine Learning · Computer Science 2023-11-28 Daniel Gomes de Pinho Zanco , Leszek Szczecinski , Eduardo Vinicius Kuhn , Rui Seara

This work is concerned with the rating of players/teams in face-to-face games with three possible outcomes: loss, win, and draw. This is one of the fundamental problems in sport analytics, where the very simple and popular, non-trivial…

Statistics Theory · Mathematics 2019-10-15 Leszek Szczecinski , Aymen Djebbi
‹ Prev 1 2 3 10 Next ›