中文
相关论文

相关论文: Using Elo Rating as a Metric for Comparative Judge…

200 篇论文

How to better reduce measurement variability and bias introduced by subjectivity in crowdsourced labelling remains an open question. We introduce a theoretical framework for understanding how random error and measurement bias enter into…

人机交互 · 计算机科学 2023-12-05 Hasti Narimanzadeh , Arash Badie-Modiri , Iuliia Smirnova , Ted Hsuan Yun Chen

Rating systems play a crucial role in evaluating player skill across competitive environments. The Elo rating system, originally designed for deterministic and information-complete games such as chess, has been widely adopted and modified…

计算机科学与博弈论 · 计算机科学 2025-12-23 Avirup Chakraborty , Shirsa Maitra , Tathagata Banerjee , Diganta Mukherjee , Tridib Mukherjee

This paper introduces a score-driven rating system, a generalization of the classical Elo rating system that employs the score, i.e. the gradient of the log-likelihood, as the updating mechanism for player and team ratings. The proposed…

机器学习 · 计算机科学 2026-04-13 Vladimír Holý , Michal Černý

The Elo rating system has been used world wide for individual sports and team sports, as exemplified by the European Go Federation (EGF), International Chess Federation (FIDE), International Federation of Association Football (FIFA), and…

人工智能 · 计算机科学 2021-05-04 Ben Wise

Many environments assign several Elo ratings to the same agent: a chess player has classical, rapid, and blitz ratings; an online platform may rate by time control, mode, or format; an evaluator may rate performance across tasks or roles.…

理论经济学 · 经济学 2026-05-12 Mehmet Mars Seven

As Large Language Models (LLMs) achieve breakthroughs in complex reasoning, Codeforces-based Elo ratings have emerged as a prominent metric for evaluating competitive programming capabilities. However, these ratings are often reported…

In arena-style evaluation of large language models (LLMs), two LLMs respond to a user query, and the user chooses the winning response or deems the "battle" a draw, resulting in an adjustment to the ratings of both models. The prevailing…

计算与语言 · 计算机科学 2025-10-03 Raphael Tang , Crystina Zhang , Wenyan Li , Carmen Lai , Pontus Stenetorp , Yao Lu

The Elo rating system is a popular and widely adopted method for measuring the relative skill levels of players or teams in various sports and competitions. It assigns players numerical ratings and dynamically updates them based on game…

概率论 · 数学 2026-01-28 Roberto Cortez , Hagop Tossounian

The TextClass Benchmark project is an ongoing, continuous benchmarking process that aims to provide a comprehensive, fair, and dynamic evaluation of LLMs and transformers for text classification tasks. This evaluation spans various domains…

计算与语言 · 计算机科学 2024-12-10 Bastián González-Bustamante

In competitive games, strength ratings like Elo are widely used to quantify player skill and support matchmaking by accounting for skill disparities better than simple win rate statistics. However, scalar ratings cannot handle complex…

机器学习 · 计算机科学 2025-02-07 Chiu-Chou Lin , I-Chen Wu

We suggest an improvement of the Elo rating system. Whereas Elo's theoretical background remains unaffected, we significantly change the way in which rating values are adjusted. It turns out that the modified system behaves much more…

经典分析与常微分方程 · 数学 2018-01-17 Fabian Langholf

This study aims to provide a data-driven approach for empirically tuning and validating rating systems, focusing on the Elo system. Well-known rating frameworks, such as Elo, Glicko, TrueSkill systems, rely on parameters that are usually…

应用统计 · 统计学 2025-12-23 Shirsa Maitra , Tathagata Banerjee , Anushka De , Diganta Mukherjee , Tridib Mukherjee

The Elo algorithm, renowned for its simplicity, is widely used for rating in sports tournaments and other applications. However, despite its widespread use, a detailed understanding of the convergence characteristics of the Elo algorithm is…

机器学习 · 计算机科学 2023-11-28 Daniel Gomes de Pinho Zanco , Leszek Szczecinski , Eduardo Vinicius Kuhn , Rui Seara

Large Language Models (LLMs) have shown promise in Automated Essay Scoring (AES), but their zero-shot and few-shot performance often falls short compared to state-of-the-art models and human raters. However, fine-tuning LLMs for each…

计算与语言 · 计算机科学 2024-07-09 Seungju Kim , Meounggun Jo

Evaluation has traditionally focused on ranking candidates for a specific skill. Modern generalist models, such as Large Language Models (LLMs), decidedly outpace this paradigm. Open-ended evaluation systems, where candidate models are…

计算机科学与博弈论 · 计算机科学 2025-05-09 Siqi Liu , Ian Gemp , Luke Marris , Georgios Piliouras , Nicolas Heess , Marc Lanctot

We introduce a novel system of matching and scoring players in tournaments, called Multi-Tier Tournaments, illustrated by chess and based on the following rules: 1. Players are divided into skill-based tiers, based on their Elo ratings. 2.…

理论经济学 · 经济学 2024-07-22 Steven J. Brams , Mehmet S. Ismail

As intelligent agents become more generally-capable, i.e. able to master a wide variety of tasks, the complexity and cost of properly evaluating them rises significantly. Tasks that assess specific capabilities of the agents can be…

人工智能 · 计算机科学 2026-02-12 Marc Lanctot , Kate Larson , Ian Gemp , Michael Kaisers

Accurate estimation of question difficulty and prediction of student performance play key roles in optimizing educational instruction and enhancing learning outcomes within digital learning platforms. The Elo rating system is widely…

计算机与社会 · 计算机科学 2024-03-14 Erva Nihan Kandemir , Jill-Jenn Vie , Adam Sanchez-Ayte , Olivier Palombi , Franck Ramus

This paper investigates the evaluation of learned multiagent strategies in the incomplete information setting, which plays a critical role in ranking and training of agents. Traditionally, researchers have relied on Elo ratings for this…

多智能体系统 · 计算机科学 2020-01-13 Mark Rowland , Shayegan Omidshafiei , Karl Tuyls , Julien Perolat , Michal Valko , Georgios Piliouras , Remi Munos

It was recently observed that Elo ratings fail at preserving transitive relations among strategies and therefore cannot correctly extract the transitive component of a game. We provide a characterization of transitive games as a weak…

计算机科学与博弈论 · 计算机科学 2024-03-07 Nelson Vadori , Rahul Savani