English
Related papers

Related papers: Do peers share the same criteria for assessing gra…

200 papers

Researchers addressing post-treatment complications in randomized trials often turn to principal stratification to define relevant assumptions and quantities of interest. One approach for estimating causal effects in this framework is to…

Methodology · Statistics 2016-06-09 Avi Feller , Fabrizia Mealli , Luke Miratrix

Online platforms for social interactions are an essential part of modern society. With the advance of technology and the rise of algorithms and AI, content is now filtered systematically and facilitates the formation of filter bubbles. This…

Crowdsourcing offers an affordable and scalable means to collect relevance judgments for IR test collections. However, crowd assessors may show higher variance in judgment quality than trusted assessors. In this paper, we investigate how to…

Information Retrieval · Computer Science 2018-06-12 Mucahid Kutlu , Tyler McDonnell , Aashish Sheshadri , Tamer Elsayed , Matthew Lease

Double-blind peer review mechanism has become the skeleton of academic research across multiple disciplines including computer science, yet several studies have questioned the quality of peer reviews and raised concerns on potential biases…

Computers and Society · Computer Science 2022-11-14 Jiayao Zhang , Hongming Zhang , Zhun Deng , Dan Roth

Disagreements are common in online discussions. Disagreement may foster collaboration and improve the quality of a discussion under some conditions. Although there exist methods for recognizing disagreement, a deeper understanding of…

Computation and Language · Computer Science 2024-10-03 Michiel van der Meer , Piek Vossen , Catholijn M. Jonker , Pradeep K. Murukannaiah

Instead of testing for unanimous agreement, I propose learning how broad of a consensus favors one distribution over another (of earnings, productivity, asset returns, test scores, etc.). Specifically, given a sample from each of two…

Econometrics · Economics 2024-08-27 David M. Kaplan

A semi-supervised model of peer review is introduced that is intended to overcome the bias and incompleteness of traditional peer review. Traditional approaches are reliant on human biases, while consensus decision-making is constrained by…

Digital Libraries · Computer Science 2013-11-12 Bradly Alicea

Finding the right reviewers to assess the quality of conference submissions is a time consuming process for conference organizers. Given the importance of this step, various automated reviewer-paper matching solutions have been proposed to…

Computation and Language · Computer Science 2019-09-26 Omer Anjum , Hongyu Gong , Suma Bhat , Wen-Mei Hwu , Jinjun Xiong

The National Science Foundation (NSF) will be experimenting with a new distributed approach to reviewing proposals, whereby a group of principal investigators (PIs) or proposers in a subfield act as reviewers for the proposals submitted by…

Computer Science and Game Theory · Computer Science 2013-07-25 Parinaz Naghizadeh , Mingyan Liu

Humans tend to strongly agree on ratings on a scale for extreme cases (e.g., a CAT is judged as very concrete), but judgements on mid-scale words exhibit more disagreement. Yet, collected rating norms are heavily exploited across…

Computation and Language · Computer Science 2024-04-18 Urban Knupleš , Diego Frassinelli , Sabine Schulte im Walde

One of the major challenges for collective intelligence is inconsistency, which is unavoidable whenever subjective assessments are involved. Pairwise comparisons allow one to represent such subjective assessments and to process them by…

Artificial Intelligence · Computer Science 2015-08-06 J. Fueloep , W. W. Koczkodaj , S. J. Szarek

We propose the PeerRank method for peer assessment. This constructs a grade for an agent based on the grades proposed by the agents evaluating the agent. Since the grade of an agent is a measure of their ability to grade correctly, the…

Artificial Intelligence · Computer Science 2014-05-29 Toby Walsh

We present the NeurIPS 2021 consistency experiment, a larger-scale variant of the 2014 NeurIPS experiment in which 10% of conference submissions were reviewed by two independent committees to quantify the randomness in the review process.…

Machine Learning · Computer Science 2023-06-07 Alina Beygelzimer , Yann N. Dauphin , Percy Liang , Jennifer Wortman Vaughan

Peer review is vital in academia for evaluating research quality. Top AI conferences use reviewer confidence scores to ensure review reliability, but existing studies lack fine-grained analysis of text-score consistency, potentially missing…

Computation and Language · Computer Science 2025-05-22 Wenqing Wu , Haixu Xi , Chengzhi Zhang

Peer assessment is a popular technique for a more fine-grained evaluation of individual students in group projects. Its effect on the evaluation is well studied. However, its effects on the learning abilities of students are often…

Software Engineering · Computer Science 2020-12-08 Wouter Groeneveld , Joost Vennekens , Kris Aerts

This study investigates the reliability and validity of five advanced Large Language Models (LLMs), Claude 3.5, DeepSeek v2, Gemini 2.5, GPT-4, and Mistral 24B, for automated essay scoring in a real world higher education context. A total…

Computers and Society · Computer Science 2025-08-05 Andrea Gaggioli , Giuseppe Casaburi , Leonardo Ercolani , Francesco Collova' , Pietro Torre , Fabrizio Davide

In this study we present a metric of consensus for Likert scales. The measure gives the level of agreement as the percentage of consensus among respondents. The proposed framework allows to design a positional indicator that gives the…

Methodology · Statistics 2018-10-26 Oscar Claveria

A recent work of the authors on the analysis of pairwise comparison matrices that can be made consistent by the modification of a few elements is continued and extended. Inconsistency indices are defined for indicating the overall quality…

Optimization and Control · Mathematics 2015-11-05 Sándor Bozóki , János Fülöp , Attila Poesz

Given a pre-trained classifier and multiple human experts, we investigate the task of online classification where model predictions are provided for free but querying humans incurs a cost. In this practical but under-explored setting,…

Machine Learning · Computer Science 2023-12-14 Sam Showalter , Alex Boyd , Padhraic Smyth , Mark Steyvers

This study examines whether there is any evidence of bias in two areas of common critique of open, non-anonymous peer review - and used in the post-publication, peer review system operated by the open-access scholarly publishing platform…

Digital Libraries · Computer Science 2019-11-11 Mike Thelwall , Verena Weigert , Liz Allen , Zena Nyakoojo , Eleanor-Rose Papas