中文
相关论文

相关论文: Using Elo Rating as a Metric for Comparative Judge…

200 篇论文

With the rapid development of large language models (LLM), the evaluation of LLM becomes increasingly important. Measuring text generation tasks such as summarization and article creation is very difficult. Especially in specific…

计算与语言 · 计算机科学 2025-09-25 Kaiqi Zhang , Shuai Yuan , Honghan Zhao

National research evaluation exercises provide a comparative measure of research performance of the nation's institutions, and as such represent a tool for stimulating research productivity, particularly if the results are used to inform…

数字图书馆 · 计算机科学 2018-11-06 Giovanni Abramo , Ciriaco Andrea D'Angelo , Flavia Di Costa

Existing LLM-as-a-Judge approaches for evaluating text generation suffer from rating inconsistencies, with low agreement and high rating variance across different evaluator models. We attribute this to subjective evaluation criteria…

计算与语言 · 计算机科学 2025-11-04 Yukyung Lee , Joonghoon Kim , Jaehee Kim , Hyowon Cho , Jaewook Kang , Pilsung Kang , Najoung Kim

Rankings and scores are two common data types used by judges to express preferences and/or perceptions of quality in a collection of objects. Numerous models exist to study data of each type separately, but no unified statistical model…

统计方法学 · 统计学 2022-09-02 Michael Pearce , Elena A. Erosheva

The $\textit{LLM-as-a-judge}$ paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: that judges evaluate text strictly on its semantic content, impervious to surrounding…

人工智能 · 计算机科学 2026-04-17 Manan Gupta , Inderjeet Nair , Lu Wang , Dhruv Kumar

The recent proliferation of artificial intelligence and machine learning (AI/ML) systems highlights the need for all people to develop effective competencies to interact with and examine AI/ML systems. We study shifts in five experienced…

人机交互 · 计算机科学 2026-03-30 Daniel J. Noh , Deborah A. Fields , Yasmin B. Kafai , Danaé Metaxa

Large language models (LLMs) are increasingly deployed as automatic judges to evaluate system outputs in tasks such as summarization, dialogue, and creative writing. A faithful judge should base its verdicts solely on response quality and…

计算与语言 · 计算机科学 2025-10-15 Arash Marioriyad , Mohammad Hossein Rohban , Mahdieh Soleymani Baghshah

Systematic evaluations of publicly funded research typically employ a combination of bibliometrics and peer review, but it is not known whether the bibliometric component introduces biases. This article compares three alternative mechanisms…

数字图书馆 · 计算机科学 2022-12-16 Mike Thelwall , Kayvan Kousha , Mahshid Abdoli , Emma Stuart , Meiko Makita , Paul Wilson , Jonathan Levitt

The rapid adoption of large language models in AI-powered language education has created an urgent need for evaluations that assess pedagogical effectiveness, particularly in language learning--one of the most common LLM use cases (Tamkin…

计算机与社会 · 计算机科学 2026-05-25 James Edgell , Wm. Matthew Kennedy , Isaac Pattis , Ben Knight , Danielle Carvalho , Elizabeth Wonnacott

Pragmatic reasoning, inferring intended meaning beyond literal semantics, underpins everyday communication yet remains difficult for large language models. We present the Contextual Emotional Inference (CEI) Benchmark: 300 human-validated…

Classifiers commonly make use of pre-annotated datasets, wherein a model is evaluated by pre-defined metrics on a held-out test set typically made of human-annotated labels. Metrics used in these evaluations are tied to the availability of…

计算与语言 · 计算机科学 2021-06-15 Yifan Ding , Nicholas Botzer , Tim Weninger

In line with the principle of honesty, there has been a growing effort to train large language models (LLMs) to generate outputs containing epistemic markers. However, evaluation in the presence of epistemic markers has been largely…

计算与语言 · 计算机科学 2025-05-02 Dongryeol Lee , Yerin Hwang , Yongil Kim , Joonsuk Park , Kyomin Jung

To overcome the limitations of automated metrics (e.g. BLEU, METEOR) for evaluating dialogue systems, researchers typically use human judgments to provide convergent evidence. While it has been demonstrated that human judgments can suffer…

计算与语言 · 计算机科学 2019-09-24 Sashank Santhanam , Samira Shaikh

Equity of educational outcome and fairness of AI with respect to race have been topics of increasing importance in education. In this work, we address both with empirical evaluations of grade prediction in higher education, an important…

计算机与社会 · 计算机科学 2021-05-17 Weijie Jiang , Zachary A. Pardos

Given the rapid progress of generative AI, there is a pressing need to systematically compare and choose between the numerous models and configurations available. The scale and versatility of such evaluations make the use of LLM-based…

计算与语言 · 计算机科学 2025-06-11 Ariel Gera , Odellia Boni , Yotam Perlitz , Roy Bar-Haim , Lilach Eden , Asaf Yehudai

Understanding the correlation between two different scores for the same set of items is a common problem in information retrieval, and the most commonly used statistics that quantifies this correlation is Kendall's $\tau$. However, the…

社会与信息网络 · 计算机科学 2014-11-03 Sebastiano Vigna

Long chain-of-thought (CoT) trajectories provide rich supervision signals for distilling reasoning from teacher to student LLMs. However, both prior work and our experiments show that trajectories from stronger teachers do not necessarily…

Detailed feedback on courses and lecture content is essential for their improvement and also serves as a tool for reflection. However, feedback methods are often only used sporadically, especially in mass courses, because collecting and…

计算机与社会 · 计算机科学 2024-04-16 Armin Egetenmeier , Sven Strickroth

A human-like chess engine should mimic the style, errors, and consistency of a strong human player rather than maximize playing strength. We show that training from move sequences alone forces a model to learn two capabilities: state…

人工智能 · 计算机科学 2026-04-01 Quanhao Li , Wei Jiang

Peer assessment is an efficient and effective learning assessment method that has been used widely in diverse fields in higher education. Despite its many benefits, a fundamental problem in peer assessment is that participants lack the…

计算机与社会 · 计算机科学 2015-06-19 Yanqing Wang , Yaowen Liang , Luning Liu , Ying Liu