English
Related papers

Related papers: An Augmented Rating System for Test cricket: adapt…

200 papers

The Elo system for rating chess players, also used in other games and sports, was adopted by the World Chess Federation over four decades ago. Although not without controversy, it is accepted as generally reliable and provides a method for…

Physics and Society · Physics 2011-03-31 Trevor Fenner , Mark Levene , George Loizou

In Natural Language Processing (NLP), the Elo rating system, originally designed for ranking players in dynamic games such as chess, is increasingly being used to evaluate Large Language Models (LLMs) through "A vs B" paired comparisons.…

Computation and Language · Computer Science 2023-11-30 Meriem Boubdir , Edward Kim , Beyza Ermis , Sara Hooker , Marzieh Fadaee

Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmaxxing or training on the test task. In this work, we initiate…

Computer Science and Game Theory · Computer Science 2026-03-10 Yatong Chen , Guanhua Zhang , Moritz Hardt

In this manuscript, we concentrate on a specific type of covariates, which we call statistically enhanced, for modeling tennis matches for men at Grand slam tournaments. Our goal is to assess whether these enhanced covariates have the…

Applications · Statistics 2025-02-27 Nourah Buhamra , Andreas Groll

As the technology advances, an ample amount of data is collected in sports with the help of advanced sensors. Sports Analytics is the study of this data to provide a constructive advantage to the team and its players. The game of…

Machine Learning · Computer Science 2020-11-05 Rahul Chakwate , Madhan R A

Despite the availability of benchmark machine learning (ML) repositories (e.g., UCI, OpenML), there is no standard evaluation strategy yet capable of pointing out which is the best set of datasets to serve as gold standard to test different…

Ranking a vector of alternatives on the basis of a series of paired comparisons is a relevant topic in many instances. A popular example is ranking contestants in sport tournaments. To this purpose, paired comparison models such as the…

Applications · Statistics 2013-01-15 Guido Masarotto , Cristiano Varin

We present a Bayesian rating system based on the method of paired comparisons. Our system is a flexible generalization of the well-known Glicko, and in particular can better accommodate games with significant elements of luck. Our system is…

Methodology · Statistics 2023-03-28 Alex Cowan

We propose a player rating mechanism for Counter-Strike: Global Offensive (CS ), a popular e-sport, by analyzing players' Plus/Minus values. The Plus/Minus value represents the average point difference between a player's team and the…

Applications · Statistics 2024-09-10 Hongyu Xu , Sarat Moka

The problem of testing the reliability of ensemble forecasting systems is revisited. A popular tool to assess the reliability of ensemble forecasting systems (for scalar verifications) is the rank histogram, this histogram is expected to be…

Atmospheric and Oceanic Physics · Physics 2018-12-26 Jochen Bröcker

In the sport of cricket, player batting ability is traditionally measured using the batting average. However, the batting average fails to measure both short-term changes in ability that occur during an innings, and long-term changes that…

Applications · Statistics 2021-03-25 Oliver George Stevenson , Brendon James Brewer

Artificial intelligence-based systems for player risk detection have become central to harm prevention efforts in the gambling industry. However, growing concerns around transparency and effectiveness have highlighted the absence of…

Robust estimation is primarily concerned with providing reliable parameter estimates in the presence of outliers. Numerous robust loss functions have been proposed in regression and classification, along with various computing algorithms.…

Methodology · Statistics 2024-02-26 Zhu Wang

Even though Google Research Football (GRF) was initially benchmarked and studied as a single-agent environment in its original paper, recent years have witnessed an increasing focus on its multi-agent nature by researchers utilizing it as a…

Multiagent Systems · Computer Science 2023-09-25 Yan Song , He Jiang , Haifeng Zhang , Zheng Tian , Weinan Zhang , Jun Wang

Confronted with the challenge of identifying the most suitable metric to validate the merits of newly proposed models, the decision-making process is anything but straightforward. Given that comparing rankings introduces its own set of…

Information Retrieval · Computer Science 2024-08-30 Chiara Balestra , Andreas Mayr , Emmanuel Müller

In this work, we deal with the problem of rating in sports, where the skills of the players/teams are inferred from the observed outcomes of the games. Our focus is on the online rating algorithms which estimate the skills after each new…

Machine Learning · Statistics 2021-04-30 Leszek Szczecinski , Raphaëlle Tihon

We analyze cross-correlation between runs scored over a time interval in cricket matches of different teams using methods of random matrix theory (RMT). We obtain an ensemble of cross-correlation matrices $C$ from runs scored by eight…

Applications · Statistics 2015-02-12 Manu Kalia , Saugata Ghosh

The Elo rating system, which was originally proposed by Arpad Elo for chess, has become one of the most important rating systems in sports, economics and gaming nowadays. Its original formulation is based on two-player zero-sum games, but…

Optimization and Control · Mathematics 2022-04-12 Düring Bertram , Fischer Michael , Wolfram Marie-Therese

This paper proposes a multiple-membership generalized linear mixed model for ranking college football teams using only their win/loss records. The model results in an intractable, high-dimensional integral due to the random effects…

Applications · Statistics 2014-04-01 Andrew T. Karl

Models that top leaderboards often perform unsatisfactorily when deployed in real world applications; this has necessitated rigorous and expensive pre-deployment model testing. A hitherto unexplored facet of model performance is: Are our…

Computation and Language · Computer Science 2021-06-11 Swaroop Mishra , Anjana Arunkumar