中文
相关论文

相关论文: Judging the Judges: Evaluating the Performance of …

200 篇论文

The monitoring of judges and referees in sports has become an important topic due to the increasing media exposure of international sporting events and the large monetary sums involved. In this article, we present a method to assess the…

应用统计 · 统计学 2019-08-20 Sandro Heiniger , Hugues Mercier

National bias in sports judging is a well-known issue and has been observed in several sports: judges, in the aggregate, give higher marks to athletes of the same nationality. In this work, we study the national bias of international…

应用统计 · 统计学 2019-08-19 Sandro Heiniger , Hugues Mercier

Functional fitness movements are widely used in training, competition, and health-oriented exercise programs, yet consistently enforcing repetition (rep) standards remains challenging due to subjective human judgment, time constraints, and…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Shaibal Saha , Fan Li , Yunge Li , Arun Iyengar , Lucas Alves , Lanyu Xu

Scaling test-time computation, or affording a generator large language model (LLM) extra compute during inference, typically employs the help of external non-generative evaluators (i.e., reward models). Concurrently, LLM-judges, models…

计算与语言 · 计算机科学 2025-05-23 Yilun Zhou , Austin Xu , Peifeng Wang , Caiming Xiong , Shafiq Joty

To reduce the need for human annotations, large language models (LLMs) have been proposed as judges of the quality of other candidate models. The performance of LLM judges is typically evaluated by measuring the correlation with human…

计算与语言 · 计算机科学 2025-05-14 Andreas Stephan , Dawei Zhu , Matthias Aßenmacher , Xiaoyu Shen , Benjamin Roth

Automatic fault detection is a major challenge in many sports. In race walking, referees visually judge faults according to the rules. Hence, ensuring objectivity and fairness while judging is important. To address this issue, some studies…

计算机视觉与模式识别 · 计算机科学 2022-08-29 Tomohiro Suzuki , Kazuya Takeda , Keisuke Fujii

Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such evaluations make the use of LLM-based…

计算与语言 · 计算机科学 2025-06-11 Ariel Gera , Odellia Boni , Yotam Perlitz , Roy Bar-Haim , Lilach Eden , Asaf Yehudai

LLM-as-judge evaluation has become standard practice for open-ended model assessment; however, judges exhibit systematic biases that cannot be averaged out by increasing the number of scenarios or generations. These biases are often similar…

计算与语言 · 计算机科学 2026-05-05 Ziyi Zhu , Olivier Tieleman , Alexey Bukhtiyarov , Jinghong Chen

AI-driven Action Quality Assessment (AQA) of sports videos can mimic Olympic judges to help score performances as a second opinion or for training. However, these AI methods are uninterpretable and do not justify their scores, which is…

人工智能 · 计算机科学 2023-03-17 Hitoshi Matsuyama , Nobuo Kawaguchi , Brian Y. Lim

Though athletics statistics are abundant, it is a difficult task to quantitatively compare performances from different events of track, field, and road running in a meaningful way. There are several commonly-used methods, but each has its…

应用统计 · 统计学 2014-08-27 Brian Godsey

A common form of competition is one where judges grade contestants' performances which are then compiled to determine the final ranking of the contestants. Unlike in another common form of competition where two contestants play a…

物理与社会 · 物理学 2016-08-09 Gyuhyeon Jeon , Juyong Park

Scoring rules are widely used to rank athletes in sports and candidates in elections. Each position in each individual ranking is worth a certain number of points; the total sum of points determines the aggregate ranking. The question is…

计算机科学与博弈论 · 计算机科学 2022-09-09 Aleksei Y. Kondratev , Egor Ianovski , Alexander S. Nesterov

Humans are routinely asked to evaluate the performance of other individuals, separating success from failure and affecting outcomes from science to education and sports. Yet, in many contexts, the metrics driving the human evaluation…

物理与社会 · 物理学 2017-12-07 Luca Pappalardo , Paolo Cintia , Dino Pedreschi , Fosca Giannotti , Albert-Laszlo Barabasi

LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized. As LLMs become more…

Instruction-following is a foundational capability of large language models (LLMs), with its improvement hinging on scalable and accurate feedback from judge models. However, the reliability of current judge models in instruction-following…

计算与语言 · 计算机科学 2026-04-17 Bosi Wen , Yilin Niu , Cunxiang Wang , Xiaoying Ling , Ying Zhang , Pei Ke , Hongning Wang , Minlie Huang

LLM-as-a-Judge has emerged as a promising alternative to human evaluators across various tasks, yet inherent biases - particularly position bias, the tendency to favor solutions based on their position within the prompt - compromise its…

计算与语言 · 计算机科学 2025-11-12 Lin Shi , Chiyu Ma , Wenhua Liang , Xingjian Diao , Weicheng Ma , Soroush Vosoughi

Instruction-based Image Editing (IIE) models have made significantly improvement due to the progress of multimodal large language models (MLLMs) and diffusion models, which can understand and reason about complex editing instructions. In…

人工智能 · 计算机科学 2025-04-11 Chenxi Sun , Hongzhi Zhang , Qi Wang , Fuzheng Zhang

Pairwise comparisons from multiple judges are central to large language model evaluation and preference modeling, yet standard ranking pipelines often pool judgments into a single score vector, treating systematic judge disagreement as…

统计方法学 · 统计学 2026-05-08 Shibo Yu , Yingzhou Wang , Yan Chen , Guodong Li , Jin-Hong Du

This study examines the role of human judges in legal decision-making by using machine learning to predict child physical custody outcomes in French appellate courts. Building on the legal realism-formalism debate, we test whether…

计算与语言 · 计算机科学 2025-07-21 Guillaume Zambrano

Large language models (LLMs) can serve as judges that offer rapid and reliable assessments of other LLM outputs. However, models may systematically assign overly favorable ratings to their own outputs, a phenomenon known as self-bias, which…

‹ 上一页 1 2 3 10 下一页 ›