中文
相关论文

相关论文: Using Elo Rating as a Metric for Comparative Judge…

200 篇论文

As the importance of comprehensive evaluation in workshop courses increases, there is a growing demand for efficient and fair assessment methods that reduce the workload for faculty members. This paper presents an evaluation conducted with…

计算机与社会 · 计算机科学 2024-05-30 Toru Ishida , Tongxi Liu , Hailong Wang , William K. Cheung

Measuring long-run LLM outcomes (user satisfaction, expert judgment, downstream KPIs) is expensive. Teams default to cheap LLM judges, but uncalibrated proxies can invert rankings entirely. Causal Judge Evaluation (CJE) makes it affordable…

统计方法学 · 统计学 2026-01-22 Eddie Landesberg , Manjari Narayan

Reinforcement Learning Fine-Tuning (RLFT) has achieved notable success in tasks with objectively verifiable answers (e.g., code generation, mathematical reasoning), yet struggles with open-ended subjective tasks like role-playing dialogue.…

计算与语言 · 计算机科学 2025-08-13 Xinge Ye , Rui Wang , Yuchuan Wu , Victor Ma , Feiteng Fang , Fei Huang , Yongbin Li

Popular metrics used for evaluating image captioning systems, such as BLEU and CIDEr, provide a single score to gauge the system's overall effectiveness. This score is often not informative enough to indicate what specific errors are made…

计算与语言 · 计算机科学 2019-09-06 Ming Jiang , Junjie Hu , Qiuyuan Huang , Lei Zhang , Jana Diesner , Jianfeng Gao

The large language model (LLM)-as-judge paradigm has been used to meet the demand for a cheap, reliable, and fast evaluation of model outputs during AI system development and post-deployment monitoring. While judge models -- LLMs finetuned…

计算与语言 · 计算机科学 2025-03-21 Austin Xu , Srijan Bansal , Yifei Ming , Semih Yavuz , Shafiq Joty

Peer review is a cornerstone of scientific publishing, including at premier machine learning conferences such as ICLR. As submission volumes increase, understanding the nature and dynamics of the review process is crucial for improving its…

计算机与社会 · 计算机科学 2025-11-20 Amir Hossein Kargaran , Nafiseh Nikeghbal , Jing Yang , Nedjma Ousidhoum

Combining short-term experimental data with observational data enables credible long-term policy evaluation. The literature offers two key but non-nested assumptions, namely the latent unconfoundedness (LU; Athey et al., 2020) and…

计量经济学 · 经济学 2024-01-23 Yechan Park , Yuya Sasaki

Many applications of AI involve scoring individuals using a learned function of their attributes. These predictive risk scores are then used to take decisions based on whether the score exceeds a certain threshold, which may vary depending…

机器学习 · 统计学 2021-02-26 Robin Vogel , Aurélien Bellet , Stephan Clémençon

Automated essay scoring (AES) research often relies on rank-based correlation metrics to validate analytic assessment. However, such metrics obscure both intrinsic intercorrelations among analytic dimensions that arise from the structure of…

计算与语言 · 计算机科学 2026-05-07 Stefano Bannò , Kate Knill , Mark Gales

Recent advances in image generation, often driven by proprietary systems like GPT-4o Image Gen, regularly introduce new capabilities that reshape how users interact with these models. Existing benchmarks often lag behind and fail to capture…

计算机视觉与模式识别 · 计算机科学 2025-10-20 Jiaxin Ge , Grace Luo , Heekyung Lee , Nishant Malpani , Long Lian , XuDong Wang , Aleksander Holynski , Trevor Darrell , Sewon Min , David M. Chan

In programming education, providing manual feedback is essential but labour-intensive, posing challenges in consistency and timeliness. We introduce ECHO, a machine learning method to automate the reuse of feedback in educational code…

This paper concerns with statistical estimation and inference for the ranking problems based on pairwise comparisons with additional covariate information such as the attributes of the compared items. Despite extensive studies, few prior…

统计方法学 · 统计学 2024-03-26 Jianqing Fan , Jikai Hou , Mengxin Yu

Learning-to-rank (LTR) algorithms are ubiquitous and necessary to explore the extensive catalogs of media providers. To avoid the user examining all the results, its preferences are used to provide a subset of relatively small size. The…

Ideal or real - that is the question.In this work, we explore whether principles from game theory can be effectively applied to the evaluation of large language models (LLMs). This inquiry is motivated by the growing inadequacy of…

计算与语言 · 计算机科学 2026-04-07 Gao Yang , Yuhang Liu , Siyu Miao , Xinyue Liang , Zhengyang Liu , Heyan Huang

Entity linking (EL) is the task of disambiguating mentions in text by associating them with entries in a predefined database of mentions (persons, organizations, etc). Most previous EL research has focused mainly on one language, English,…

计算与语言 · 计算机科学 2017-12-06 Avirup Sil , Radu Florian

One of the most popular club football tournaments, the UEFA Champions League, will see a fundamental reform from the 2024/25 season: the traditional group stage will be replaced by one league where each of the 36 teams plays eight matches.…

应用统计 · 统计学 2024-04-03 László Csató

Assessing and comparing player skill in online multiplayer gaming environments is essential for fair matchmaking and player engagement. Traditional ranking models like Elo and Glicko-2, designed for two-player games, are insufficient for…

人机交互 · 计算机科学 2024-01-12 Vivek Joshy

A method is presented for evaluating authors on the basis of citations. It assigns to each author a citation score which depends upon the number of times he is cited, and upon the scores of the citers. The scores are found to be the…

历史与综述 · 数学 2008-10-07 Joseph B. Keller

In this work, we leverage a generative data model considering comparison noise to develop a fast, precise, and informative ranking algorithm from pairwise comparisons that produces a measure of confidence on each comparison. The problem of…

机器学习 · 计算机科学 2025-07-24 Filipa Valdeira , Cláudia Soares

Iterative peer grading activities may keep students engaged during in-class project presentations. Effective methods for collecting and aggregating peer assessment data are essential. Students tend to grade projects favorably. So, while…

计算机科学与博弈论 · 计算机科学 2025-03-25 Lihi Dery