中文
相关论文

相关论文: Using Elo Rating as a Metric for Comparative Judge…

200 篇论文

Large Language Models (LLMs) have shown to be effective evaluators across various domains such as machine translations or the scientific domain. Current LLM-as-a-Judge approaches rely mostly on individual assessments or a single round of…

计算与语言 · 计算机科学 2025-07-10 Isik Baran Sandan , Tu Anh Dinh , Jan Niehues

A common form of competition is one where judges grade contestants' performances which are then compiled to determine the final ranking of the contestants. Unlike in another common form of competition where two contestants play a…

物理与社会 · 物理学 2016-08-09 Gyuhyeon Jeon , Juyong Park

Supporting equitable instruction is an important issue for teachers attending diverse STEM classrooms. Visual learning analytics along with effective student survey measures can support providing on time feedback to teachers in making…

人机交互 · 计算机科学 2024-01-17 Ali Raza , William R. Penuel , Tamara Sumner

Algorithmic decision systems are increasingly used in areas such as hiring, school admission, or loan approval. Typically, these systems rely on labeled data for training a classification model. However, in many scenarios, ground-truth…

机器学习 · 计算机科学 2021-07-19 Jakob Schoeffer , Niklas Kuehl , Isabel Valera

In this note, I introduce Estimated Performance Rating (PR$^e$), a novel system for evaluating player performance in sports and games. PR$^e$ addresses a key limitation of the Tournament Performance Rating (TPR) system, which is undefined…

理论经济学 · 经济学 2023-12-21 Mehmet S. Ismail

The "LLM-as-a-Judge" paradigm, using Large Language Models (LLMs) as automated evaluators, is pivotal to LLM development, offering scalable feedback for complex tasks. However, the reliability of these judges is compromised by various…

计算与语言 · 计算机科学 2026-05-22 Qingquan Li , Shaoyu Dou , Kailai Shao , Chao Chen , Haixiang Hu

The purpose of this study is to propose a model that predicts the social and psychological factors that affect the individuals collaborative learning outcome in group projects. The model is established on the basis of two theories, namely,…

计算机与社会 · 计算机科学 2016-10-18 Sara Taraman , Yasmin Hassan , Doaa Shawky , Ashraf H. Badawi

Peer review is the most common mechanism in place for assessing requests for resources in a large variety of scientific disciplines. One of the strongest criticisms to this paradigm is the limited reproducibility of the process, especially…

物理与社会 · 物理学 2018-07-04 Ferdinando Patat

Driving progress of AI models and agents requires comparing their performance on standardized benchmarks; for general agents, individual performances must be aggregated across a potentially wide variety of different tasks. In this paper, we…

In this manuscript, we concentrate on a specific type of covariates, which we call statistically enhanced, for modeling tennis matches for men at Grand slam tournaments. Our goal is to assess whether these enhanced covariates have the…

应用统计 · 统计学 2025-02-27 Nourah Buhamra , Andreas Groll

Pairwise evaluation of Large Language Models (LLMs) is a common paradigm, but it is prone to preference bias, where judges systematically favor certain outputs, such as their own. This bias leads to inconsistent and skewed rankings across…

人工智能 · 计算机科学 2025-11-18 Yang Zhang , Cunxiang Wang , Lindong Wu , Wenbo Yu , Yidong Wang , Guangsheng Bao , Jie Tang

Ranking is a ubiquitous method for focusing the attention of human evaluators on a manageable subset of options. Its use as part of human decision-making processes ranges from surfacing potentially relevant products on an e-commerce site to…

机器学习 · 计算机科学 2024-10-31 Richa Rastogi , Thorsten Joachims

Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs. Although both methodologies are…

计算与语言 · 计算机科学 2026-05-26 Russell Yang , Ruishi Chen , Pierce Kelaita , Riya Ranjan , Sibo Ma , Charles Dickens , Matthew Guillod , Megan Ma , Julian Nyarko

Graded labels are ubiquitous in real-world learning-to-rank applications, especially in human rated relevance data. Traditional learning-to-rank techniques aim to optimize the ranked order of documents. They typically, however, ignore…

信息检索 · 计算机科学 2023-06-21 Le Yan , Zhen Qin , Gil Shamir , Dong Lin , Xuanhui Wang , Mike Bendersky

Formative feedback is widely recognized as one of the most effective drivers of student learning, yet it remains difficult to implement equitably at scale. In large or low-resource courses, instructors often lack the time, staffing, and…

计算机与社会 · 计算机科学 2025-12-01 Chenyu Zhang , Xiaohang Luo

The rapid adoption of AI powered coding assistants like ChatGPT and other coding copilots is transforming programming education, raising questions about assessment practices, academic integrity, and skill development. As educators seek…

计算机与社会 · 计算机科学 2025-05-29 Santiago Berrezueta-Guzman , Stephan Krusche , Stefan Wagner

Interactive feedback, where feedback flows in both directions between teacher and student, is more effective than traditional one-way feedback. However, it is often too time-consuming for widespread use in educational practice. While Large…

人工智能 · 计算机科学 2024-09-12 Shengxin Hong , Chang Cai , Sixuan Du , Haiyue Feng , Siyuan Liu , Xiuyi Fan

Chess championships are often organised as a Swiss-system tournament, causing great challenges in ranking the participants due to the different strength of schedules and possible circular triads. The paper suggests that pairwise comparison…

应用统计 · 统计学 2016-11-03 Lászlo Csató

Feedback has a powerful influence on learning, but it is also expensive to provide. In large classes, it may even be impossible for instructors to provide individualized feedback. Peer assessment has received attention lately as a way of…

应用统计 · 统计学 2014-10-16 Dennis L. Sun , Naftali Harris , Guenther Walther , Michael Baiocchi

LLM-as-a-judge approaches are a practical and effective way of assessing a range of text tasks. However, when using pairwise comparisons to rank a set of candidates, the computational cost scales quadratically with the number of candidates,…

计算与语言 · 计算机科学 2024-11-13 Adian Liusie , Vatsal Raina , Yassir Fathullah , Mark Gales