中文
相关论文

相关论文: Using Elo Rating as a Metric for Comparative Judge…

200 篇论文

Label error is a ubiquitous problem in annotated data. Large amounts of label error substantially degrades the quality of deep learning models. Existing methods to tackle the label error problem largely focus on the classification task, and…

Peer assessment has been widely studied as a replacement for traditional evaluation, not only by reducing the professors' workload but mainly by benefiting students' engagement and learning. Although several works successfully validate its…

人机交互 · 计算机科学 2024-07-04 Francisco Sousa , Tomás Alves , Sandra Gama , Joaquim Jorge , Daniel Gonçalves

Ranking athletes by their performance in competitions and tournaments is common in every popular sport and has significant benefits that contribute to both the organization and strategic aspects of competitions. Although rankings are…

物理与社会 · 物理学 2025-08-28 Bogdán Asztalos , Boldizsár Balázs , Gergely Palla , Tamás Vicsek

Benchmarking is a fundamental practice in machine learning (ML) for comparing the performance of classification algorithms. However, traditional evaluation methods often overlook a critical aspect: the joint consideration of dataset…

机器学习 · 计算机科学 2025-04-15 Lucas Cardoso , Vitor Santos , José Ribeiro , Regiane Kawasaki , Ricardo Prudêncio , Ronnie Alves

We consider the problem of sequential evaluation, in which an evaluator observes candidates in a sequence and assigns scores to these candidates in an online, irrevocable fashion. Motivated by the psychology literature that has studied…

机器学习 · 统计学 2023-11-20 Jingyan Wang , Ashwin Pananjady

How likely is it that Magnus Carlsen will achieve an Elo rating of $2900$? This has been a goal of Magnus and is of great current interest to the chess community. Our paper uses probabilistic methods to address this question. The…

应用统计 · 统计学 2022-08-23 Sohan Bendre , Shiva Maharaj , Nick Polson , Vadim Sokolov

Competition is a primary driver of player satisfaction and engagement in multiplayer online games. Traditional matchmaking systems aim at creating matches involving teams of similar aggregated individual skill levels, such as Elo score or…

社会与信息网络 · 计算机科学 2020-06-25 Sofia M Nikolakaki , Ogheneovo Dibie , Ahmad Beirami , Nicholas Peterson , Navid Aghdaie , Kazi Zaman

Peer review (e.g., grading assignments in Massive Open Online Courses (MOOCs), academic paper review) is an effective and scalable method to evaluate the products (e.g., assignments, papers) of a large number of agents when the number of…

计算机科学与博弈论 · 计算机科学 2014-11-11 Yuanzhang Xiao , Florian Dörfler , Mihaela van der Schaar

In this work, we explore the Large Language Model (LLM) agent reviewer dynamics in an Elo-ranked review system using real-world conference paper submissions. Multiple LLM agent reviewers with different personas are engage in multi round…

计算与语言 · 计算机科学 2026-01-14 Hsiang-Wei Huang , Junbin Lu , Kuang-Ming Chen , Jenq-Neng Hwang

Ranking a vector of alternatives on the basis of a series of paired comparisons is a relevant topic in many instances. A popular example is ranking contestants in sport tournaments. To this purpose, paired comparison models such as the…

应用统计 · 统计学 2013-01-15 Guido Masarotto , Cristiano Varin

This paper presents an Elo-based rating system for programming contests, specifically Topcoder's Single Round Matches (SRMs). We introduce a logarithmic rank-based performance metric that allows single-round, multi-player contest results to…

多智能体系统 · 计算机科学 2026-02-24 Fred Batty

Teaching and Learning process of an educational institution needs to be monitored and effectively analysed for enhancement. Teaching and Learning is a vital element for an educational institution. It is also one of the criteria set by…

系统与控制 · 计算机科学 2017-06-13 Ms. Ganesan Kavitha , Dr. Lawrance Raj

In this work we develop a new algorithm for rating of teams (or players) in one-on-one games by exploiting the observed difference of the game-points (such as goals), also known as a margin of victory (MOV). Our objective is to obtain the…

统计方法学 · 统计学 2022-02-09 Leszek Szczecinski

In the computer vision and machine learning communities, as well as in many other research domains, rigorous evaluation of any new method, including classifiers, is essential. One key component of the evaluation process is the ability to…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Sébastien Piérard , Anaïs Halin , Anthony Cioppa , Adrien Deliège , Marc Van Droogenbroeck

When the college student satisfaction survey is considered in the promotion and recognition of instructors, a usual complaint is related to the impact that biased ratings have on the arithmetic mean (used as a measure of teaching…

应用统计 · 统计学 2013-02-01 Pablo Dorta-González , María Isabel Dorta-González

Online education platforms have significantly transformed the dissemination of educational resources by providing a dynamic and digital infrastructure. With the further enhancement of this transformation, the advent of Large Language Models…

人工智能 · 计算机科学 2024-09-26 Qian-Wen Zhang , Haochen Wang , Fang Li , Siyu An , Lingfeng Qiao , Liangcai Gao , Di Yin , Xing Sun

LLM-as-a-Judge evaluation has become a standard tool for assessing base model performance. However, characterizing performance via the naive estimator, i.e., raw judge outputs, is systematically biased. Recent work has proposed estimators…

机器学习 · 计算机科学 2026-05-11 James Fiedler

We propose a new framework for imitation learning -- treating imitation as a two-player ranking-based game between a policy and a reward. In this game, the reward agent learns to satisfy pairwise performance rankings between behaviors,…

机器学习 · 计算机科学 2023-01-18 Harshit Sikchi , Akanksha Saran , Wonjoon Goo , Scott Niekum

This study investigated potential scoring biases and disparities toward English Language Learners (ELLs) when using automatic scoring systems for middle school students' written responses to science assessments. We specifically focus on…

计算与语言 · 计算机科学 2025-05-21 Shuchen Guo , Yun Wang , Jichao Yu , Xuansheng Wu , Bilgehan Ayik , Field M. Watts , Ehsan Latif , Ninghao Liu , Lei Liu , Xiaoming Zhai

Kendall's tau is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations. Its focus on pairwise score comparisons is intuitive but raises the question of how ties should be…

计算与语言 · 计算机科学 2023-10-18 Daniel Deutsch , George Foster , Markus Freitag