English
Related papers

Related papers: On Testing for Biases in Peer Review

200 papers

As LLMs are increasingly applied in socially impactful settings, concerns about gender bias have prompted growing efforts both to measure and mitigate such bias. These efforts often rely on evaluation tasks that differ from natural language…

Computation and Language · Computer Science 2025-09-11 Bufan Gao , Elisa Kreiss

With the enourmous popularity of large language models, many researchers have raised ethical concerns regarding social biases incorporated in such models. Several methods to measure social bias have been introduced, but apparently these…

Computation and Language · Computer Science 2024-09-13 Sarah Schröder , Alexander Schulz , Barbara Hammer

Estimating the effects of interventions in networks is complicated when the units are interacting, such that the outcomes for one unit may depend on the treatment assignment and behavior of many or all other units (i.e., there is…

Methodology · Statistics 2014-08-15 Dean Eckles , Brian Karrer , Johan Ugander

Peer assessment has established itself as a critical pedagogical tool in academic settings, offering students timely, high-quality feedback to enhance learning outcomes. However, the efficacy of this approach depends on two factors: (1) the…

Computers and Society · Computer Science 2025-08-26 Uchswas Paul , Shail Shah , Sri Vaishnavi Mylavarapu , M. Parvez Rashid , Edward Gehringer

We study a two-institution stable matching model in which candidates from two distinct groups are evaluated using partially correlated signals that are group-biased. This extends prior work (which assumes institutions evaluate candidates in…

Physics and Society · Physics 2025-11-11 Amit Kumar , Nisheeth K. Vishnoi

Large Language Models are versatile general-task solvers, and their capabilities can truly assist people with scholarly peer review as \textit{pre-review} agents, if not as fully autonomous \textit{peer-review} agents. While incredibly…

Digital Libraries · Computer Science 2025-12-30 Akhil Pandey Akella , Harish Varma Siravuri , Shaurya Rohatgi

We present a large-scale evaluation of 30 cognitive biases in 20 state-of-the-art large language models (LLMs) under various decision-making scenarios. Our contributions include a novel general-purpose test framework for reliable and…

Computation and Language · Computer Science 2025-11-04 Simon Malberg , Roman Poletukhin , Carolin M. Schuster , Georg Groh

We present the NeurIPS 2021 consistency experiment, a larger-scale variant of the 2014 NeurIPS experiment in which 10% of conference submissions were reviewed by two independent committees to quantify the randomness in the review process.…

Machine Learning · Computer Science 2023-06-07 Alina Beygelzimer , Yann N. Dauphin , Percy Liang , Jennifer Wortman Vaughan

Bias in large language models (LLMs) has many forms, from overt discrimination to implicit stereotypes. Counterfactual bias evaluation is a widely used approach to quantifying bias and often relies on template-based probes that explicitly…

Computation and Language · Computer Science 2026-01-15 Farnaz Kohankhaki , D. B. Emerson , Jacob-Junqi Tian , Laleh Seyyed-Kalantari , Faiza Khan Khattak

Large Language Models (LLMs) are increasingly deployed in high-stakes contexts where their outputs influence real-world decisions. However, evaluating bias in LLM outputs remains methodologically challenging due to sensitivity to prompt…

Computation and Language · Computer Science 2026-01-13 William Guey , Wei Zhang , Pei-Luen Patrick Rau , Pierrick Bougault , Vitor D. de Moura , Bertan Ucar , Jose O. Gomes

Algorithmic recommendation based on noisy preference measurement is prevalent in recommendation systems. This paper discusses the consequences of such recommendation on market concentration and inequality. Binary types denoting a…

Theoretical Economics · Economics 2025-10-21 Andreas Haupt

This article is a response to an off-the-record discussion that I had at an international meeting of epidemiologists. It centered on a concern, perhaps widely spread, that measurement error adjustment methods can induce positive bias in…

Applications · Statistics 2009-02-10 Igor Burstyn

We consider large-scale studies in which thousands of significance tests are performed simultaneously. In some of these studies, the multiple testing procedure can be severely biased by latent confounding factors such as batch effects and…

Methodology · Statistics 2016-06-21 Jingshu Wang , Qingyuan Zhao , Trevor Hastie , Art B. Owen

Standard benchmarks of bias and fairness in large language models (LLMs) measure the association between the user attributes stated or implied by a prompt and the LLM's short text response, but human-AI interaction increasingly requires…

Computation and Language · Computer Science 2025-06-06 Kristian Lum , Jacy Reese Anthis , Kevin Robinson , Chirag Nagpal , Alexander D'Amour

Large Language Models (LLMs) have made substantial progress in the past several months, shattering state-of-the-art benchmarks in many domains. This paper investigates LLMs' behavior with respect to gender stereotypes, a known issue for…

Computation and Language · Computer Science 2023-08-30 Hadas Kotek , Rikker Dockum , David Q. Sun

In many large multiple testing problems the hypotheses are divided into families. Given the data, families with evidence for true discoveries are selected, and hypotheses within them are tested. Neither controlling the error-rate in each…

Statistics Theory · Mathematics 2011-06-21 Yoav Benjamini , Marina Bogomolov

Studies accumulate over time and meta-analyses are mainly retrospective. These two characteristics introduce dependencies between the analysis time, at which a series of studies is up for meta-analysis, and results within the series.…

Methodology · Statistics 2019-10-21 Judith ter Schure , Peter D. Grünwald

While the impact of social biases in language models has been recognized, prior methods for bias evaluation have been limited to binary association tests on small datasets, limiting our understanding of bias complexities. This paper…

Computation and Language · Computer Science 2025-05-27 Marta Marchiori Manerba , Karolina Stańczak , Riccardo Guidotti , Isabelle Augenstein

Despite frequent double-blind review, systemic biases related to author demographics still disadvantage underrepresented groups. We start from a simple hypothesis: if a post-review recommender is trained with an explicit fairness…

Machine Learning · Computer Science 2026-03-03 Uttamasha Anjally Oyshi , Susan Gauch

Automatic reviewing helps handle a large volume of papers, provides early feedback and quality control, reduces bias, and allows the analysis of trends. We evaluate the alignment of automatic paper reviews with human reviews using an arena…