中文
相关论文

相关论文: How do referees integrate evaluation criteria into…

200 篇论文

Criteria are an essential component of any procedure for assessing merit. Yet, little is known about the criteria peers use in assessing grant applications. In this systematic review we therefore identify and synthesize studies that examine…

计算机与社会 · 计算机科学 2020-03-13 Sven E. Hug , Mirjam Aeschbach

Peer grading has emerged as a scalable solution for assessment in large and online classrooms, offering both logistical efficiency and pedagogical value. However, designing effective peer-grading systems remains challenging due to…

计算机与社会 · 计算机科学 2025-12-02 Uchswas Paul , Ananya Mantravadi , Jash Shah , Shail Shah , Sri Vaishnavi Mylavarapu , M Parvez Rashid , Edward Gehringer

Peer assessment has established itself as a critical pedagogical tool in academic settings, offering students timely, high-quality feedback to enhance learning outcomes. However, the efficacy of this approach depends on two factors: (1) the…

计算机与社会 · 计算机科学 2025-08-26 Uchswas Paul , Shail Shah , Sri Vaishnavi Mylavarapu , M. Parvez Rashid , Edward Gehringer

The "LLM-as-a-Judge" paradigm, using Large Language Models (LLMs) as automated evaluators, is pivotal to LLM development, offering scalable feedback for complex tasks. However, the reliability of these judges is compromised by various…

计算与语言 · 计算机科学 2026-05-22 Qingquan Li , Shaoyu Dou , Kailai Shao , Chao Chen , Haixiang Hu

Peer review (e.g., grading assignments in Massive Open Online Courses (MOOCs), academic paper review) is an effective and scalable method to evaluate the products (e.g., assignments, papers) of a large number of agents when the number of…

计算机科学与博弈论 · 计算机科学 2014-11-11 Yuanzhang Xiao , Florian Dörfler , Mihaela van der Schaar

Large language models (LLMs) are now widely used to evaluate the quality of text, a field commonly referred to as LLM-as-a-judge. While prior works mainly focus on point-wise and pair-wise evaluation paradigms. Rubric-based evaluation,…

计算与语言 · 计算机科学 2026-02-03 Yuzheng Xu , Tosho Hirasawa , Tadashi Kozuno , Yoshitaka Ushiku

In this paper, we present a mathematical model to capture various factors which may influence the accuracy of a competitive group recommendation system. We apply this model to peer review systems, i.e., conference or research grants review,…

信息检索 · 计算机科学 2012-04-13 Hong Xie , John C. S. Lui

While bibliometrics are widely used for research evaluation purposes, a common theoretical framework for conceptually understanding, empirically studying, and effectively teaching its usage is lacking. In this paper, we outline such a…

数字图书馆 · 计算机科学 2019-06-26 Lutz Bornmann , Julian N. Marewski

The monitoring of judges and referees in sports has become an important topic due to the increasing media exposure of international sporting events and the large monetary sums involved. In this article, we present a method to assess the…

应用统计 · 统计学 2019-08-20 Sandro Heiniger , Hugues Mercier

Reciprocal recommender systems~(RRS), conducting bilateral recommendations between two involved parties, have gained increasing attention for enhancing matching efficiency. However, the majority of existing methods in the literature still…

信息检索 · 计算机科学 2024-08-20 Chen Yang , Sunhao Dai , Yupeng Hou , Wayne Xin Zhao , Jun Xu , Yang Song , Hengshu Zhu

In peer review, reviewers are usually asked to provide scores for the papers. The scores are then used by Area Chairs or Program Chairs in various ways in the decision-making process. The scores are usually elicited in a quantized form to…

信息检索 · 计算机科学 2022-04-13 Yusha Liu , Yichong Xu , Nihar B. Shah , Aarti Singh

Peer grading systems aggregate noisy reports from multiple students to approximate a true grade as closely as possible. Most current systems either take the mean or median of reported grades; others aim to estimate students' grading…

人工智能 · 计算机科学 2022-12-05 Hedayat Zarkoob , Greg d'Eon , Lena Podina , Kevin Leyton-Brown

Rubric-based text evaluation increasingly uses large language models (LLMs) as scalable judges, but aligning frozen black-box models with human scoring standards remains challenging. We formulate this challenge as a criteria-transfer…

计算与语言 · 计算机科学 2026-05-29 Yihan Hong , Huaiyuan Yao , Bolin Shen , Wanpeng Xu , Hua Wei , Yushun Dong

The advancement of various fields of science depends on the actions of individual scientists via the peer review process. The referees' work patterns and stochastic nature of decision making both relate to the particular features of…

物理与社会 · 物理学 2013-04-19 T. Hartonen , M. J. Alava

Items from a database are often ranked based on a combination of multiple criteria. A user may have the flexibility to accept combinations that weigh these criteria differently, within limits. On the other hand, this choice of weights can…

数据库 · 计算机科学 2023-04-27 Abolfazl Asudeh , H. V. Jagadish , Julia Stoyanovich , Gautam Das

Systematic evaluations of publicly funded research typically employ a combination of bibliometrics and peer review, but it is not known whether the bibliometric component introduces biases. This article compares three alternative mechanisms…

数字图书馆 · 计算机科学 2022-12-16 Mike Thelwall , Kayvan Kousha , Mahshid Abdoli , Emma Stuart , Meiko Makita , Paul Wilson , Jonathan Levitt

One of the virtues of peer review is that it provides a self-regulating selection mechanism for scientific work, papers and projects. Peer review as a selection mechanism is hard to evaluate in terms of its efficiency. Serious efforts to…

物理与社会 · 物理学 2015-05-19 Stefan Thurner , Rudolf Hanel

We propose the PeerRank method for peer assessment. This constructs a grade for an agent based on the grades proposed by the agents evaluating the agent. Since the grade of an agent is a measure of their ability to grade correctly, the…

人工智能 · 计算机科学 2014-05-29 Toby Walsh

Large language models (LLMs) are increasingly used as evaluators for natural language generation, applying human-defined rubrics to assess system outputs. However, human rubrics are often static and misaligned with how models internally…

计算与语言 · 计算机科学 2026-02-10 Clemencia Siro , Pourya Aliannejadi , Mohammad Aliannejadi

When predicting future events, it is common to issue forecasts that are probabilistic, in the form of probability distributions over the range of possible outcomes. Such forecasts can be evaluated using proper scoring rules. Proper scoring…

统计计算 · 统计学 2023-05-15 Sam Allen
‹ 上一页 1 2 3 10 下一页 ›