English
Related papers

Related papers: Time to adjust: Improving replicability in experim…

200 papers

Large-scale hypothesis testing is central to modern science, where controlling the False Discovery Rate (FDR) has become the standard approach to managing false positives across many simultaneous tests. Hypotheses rarely exist in isolation;…

Methodology · Statistics 2026-05-19 Binyamin Perets , Shie Mannor

Most scientific disciplines use significance testing to draw conclusions about experimental or observational data. This classical approach provides a theoretical guarantee for controlling the number of false positives across a set of…

Applications · Statistics 2023-03-06 Stanley E. Lazic

The increasing ease of data capture and storage has led to a corresponding increase in the choice of data, the type of analysis performed on that data, and the complexity of the analysis performed. The main contribution of this paper is to…

Applications · Statistics 2018-03-14 David Kohn , Nick Glozier , Ian B. Hickie , Hugh Durrant-Whyte , Sally Cripps

Artificial Intelligence (AI) is increasingly being integrated into scientific research, particularly in the social sciences, where understanding human behavior is critical. Large Language Models (LLMs) have shown promise in replicating…

Computation and Language · Computer Science 2025-06-23 Ziyan Cui , Ning Li , Huaikang Zhou

This work introduces a companion reproducible paper with the aim of allowing the exact replication of the methods, experiments, and results discussed in a previous work [5]. In that parent paper, we proposed many and varied techniques for…

Data Structures and Algorithms · Computer Science 2019-12-30 Antonio Fariña , Miguel A. Martínez-Prieto , Francisco Claude , Gonzalo Navarro , Juan J. Lastra-Díaz , Nicola Prezza , Diego Seco

Replicability issues -- referring to the difficulty or failure of independent researchers to corroborate the results of published studies -- have hindered the meaningful progression of science and eroded public trust in scientific findings.…

Large language model (LLM) agents achieve impressive single-task performance but commonly exhibit repeated failures, inefficient exploration, and limited cross-task adaptability. Existing reflective strategies (e.g., Reflexion, ReAct)…

Artificial Intelligence · Computer Science 2025-09-09 Chunlong Wu , Ye Luo , Zhibo Qu , Min Wang

In the past two decades, psychological science has experienced an unprecedented replicability crisis which uncovered several issues. Among others, statistical inference is too often viewed as an isolated procedure limited to the analysis of…

Humans must flexibly arbitrate between exploring alternatives and exploiting learned strategies, yet they frequently exhibit maladaptive persistence by continuing to execute failing strategies despite accumulating negative evidence. Here we…

Machine Learning · Computer Science 2026-03-24 Zhipeng Zhang , Hongshun He

Replicability is essential in science as it allows us to validate and verify research findings. Impagliazzo, Lei, Pitassi and Sorrell (`22) recently initiated the study of replicability in machine learning. A learning algorithm is…

Machine Learning · Computer Science 2023-04-13 Zachary Chase , Shay Moran , Amir Yehudayoff

In this paper, a new concept, i.e. ultra-chaos, is proposed for the first time. Unlike a normal-chaos, statistical properties such as the probability density functions (PDF) of an ultra-chaos are sensitive to tiny disturbances. We…

General Physics · Physics 2022-03-22 Shijun Liao , Shijie Qin

In this study, we investigate whether LLMs can be used to indicate if a study in the behavioural social sciences is replicable. Using a dataset of 14 previously replicated studies (9 successful, 5 unsuccessful), we evaluate the ability of…

Computation and Language · Computer Science 2025-03-17 Denitsa Saynova , Kajsa Hansson , Bastiaan Bruinsma , Annika Fredén , Moa Johansson

While the task of assessing the plausibility of events such as ''news is relevant'' has been addressed by a growing body of work, less attention has been paid to capturing changes in plausibility as triggered by event modification.…

Computation and Language · Computer Science 2025-07-30 Anna Golub , Beate Zywietz , Annerose Eichel

The use of machine learning (ML) methods for prediction and forecasting has become widespread across the quantitative sciences. However, there are many known methodological pitfalls, including data leakage, in ML-based science. In this…

Machine Learning · Computer Science 2022-07-15 Sayash Kapoor , Arvind Narayanan

In the context of multiple hypotheses testing, the proportion $\pi_0$ of true null hypotheses in the pool of hypotheses to test often plays a crucial role, although it is generally unknown a priori. A testing procedure using an implicit or…

Statistics Theory · Mathematics 2009-02-17 Gilles Blanchard , Etienne Roquain

Several scientific fields including psychology are undergoing a replication crisis. There are many reasons for this problem, one of which is a misuse of p-values. There are several alternatives to p-values, and in this paper we describe a…

Methodology · Statistics 2020-10-05 Brian D. Segal

Large language models have achieved remarkable success in various tasks. However, it is challenging for them to learn new tasks incrementally due to catastrophic forgetting. Existing approaches rely on experience replay, optimization…

Computation and Language · Computer Science 2026-04-15 Yukun Zhao , Lingyong Yan , Zhenyang Li , Shuaiqiang Wang , Zhumin Chen , Zhaochun Ren , Dawei Yin

Reliable human evaluation is critical to the development of successful natural language generation models, but achieving it is notoriously difficult. Stability is a crucial requirement when ranking systems by quality: consistent ranking of…

Computation and Language · Computer Science 2024-04-03 Parker Riley , Daniel Deutsch , George Foster , Viresh Ratnakar , Ali Dabirmoghaddam , Markus Freitag

Human preference evaluations are widely used to compare generative models, yet it remains unclear how many judgments are required to reliably detect small improvements. We show that when preference signal is diffuse across prompts (i.e.,…

Computation and Language · Computer Science 2026-01-16 Wilson Y. Lee

The rapid advancement of language models (LMs) necessitates robust alignment with diverse user values. However, current preference optimization approaches often fail to capture the plurality of user opinions, instead reinforcing majority…

Computation and Language · Computer Science 2024-07-25 Louis Castricato , Nathan Lile , Rafael Rafailov , Jan-Philipp Fränken , Chelsea Finn