English
Related papers

Related papers: With Little Power Comes Great Responsibility

200 papers

While running any experiment, we often have to consider the statistical power to ensure an effective study. Statistical power or power ensures that we can observe an effect with high probability if such a true effect exists. However,…

Methodology · Statistics 2023-06-21 Ajinkya K Mulay , Sean Lane , Erin Hennes

Power analyses are an important aspect of experimental design, because they help determine how experiments are implemented in practice. It is common to specify a desired level of power and compute the sample size necessary to obtain that…

Methodology · Statistics 2022-12-09 Zach Branson , Xinran Li , Peng Ding

Given the complexity of combinations of tasks, languages, and domains in natural language processing (NLP) research, it is computationally prohibitive to exhaustively test newly proposed models on each possible experimental setting. In this…

Computation and Language · Computer Science 2020-05-05 Mengzhou Xia , Antonios Anastasopoulos , Ruochen Xu , Yiming Yang , Graham Neubig

For randomized controlled trials (RCTs) with a single intervention being measured on multiple outcomes, researchers often apply a multiple testing procedure (such as Bonferroni or Benjamini-Hochberg) to adjust $p$-values. Such an adjustment…

Methodology · Statistics 2023-05-17 Kristen Hunter , Luke Miratrix , Kristin Porter

Particle physics experiments rely on the (generalised) likelihood ratio test (LRT) for searches and measurements, which consist of composite hypothesis tests. However, this test is not guaranteed to be optimal, as the Neyman-Pearson lemma…

High Energy Physics - Phenomenology · Physics 2025-11-21 James Carzon , Aishik Ghosh , Rafael Izbicki , Ann Lee , Luca Masserano , Daniel Whiteson

When designing experimental studies with human participants, experimenters must decide how many trials each participant will complete, as well as how many participants to test. Most discussion of statistical power (the ability of a study…

Neurons and Cognition · Quantitative Biology 2021-08-30 Daniel H. Baker , Greta Vilidaite , Freya A. Lygo , Anika K. Smith , Tessa R. Flack , Andre D. Gouws , Timothy J. Andrews

Two-sample network hypothesis testing is an important inference task with applications across diverse fields such as medicine, neuroscience, and sociology. Many of these testing methodologies operate under the implicit assumption that the…

Methodology · Statistics 2024-05-28 Ayushi Saxena , Vince Lyzinski

In the experimental sciences, statistical power analyses are often used before data collection to determine the required sample size. However, traditional power analyses can be costly when data are difficult or expensive to collect. We…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Peiye Zhuang , Bliss Chapman , Ran Li , Oluwasanmi Koyejo

Good (Frequentist) statistical practice requires that statistical tests be performed in order to determine if the phenomenon being observed could plausibly occur by chance if the null hypothesis is false. Good practice also requires that a…

Computers and Society · Computer Science 2024-08-06 Michael Guerzhoy

Data-driven most powerful tests are statistical hypothesis decision-making tools that deliver the greatest power against a fixed null hypothesis among all corresponding data-based tests of a given size. When the underlying data…

Statistics Theory · Mathematics 2023-03-15 Albert Vexler , Alan D. Hutson

Sample size calculations for power analysis are critical for clinical research and trial design, yet their complexity and reliance on statistical expertise create barriers for many researchers. We introduce PowerGPT, an AI-powered system…

Modern studies increasingly leverage outcomes predicted by machine learning and artificial intelligence (AI/ML) models, and recent work, such as prediction-powered inference (PPI), has developed valid downstream statistical inference…

Methodology · Statistics 2026-03-18 Yiqun T. Chen , Moran Guo , Shengy Li

Most scientific disciplines use significance testing to draw conclusions about experimental or observational data. This classical approach provides a theoretical guarantee for controlling the number of false positives across a set of…

Applications · Statistics 2023-03-06 Stanley E. Lazic

It has been demonstrated that the statistical power of many neuroscience studies is very low, so that the results are unlikely to be robustly reproducible. How are neuroscientists and the journals in which they publish responding to this…

Neurons and Cognition · Quantitative Biology 2017-01-06 Geoffrey J Goodhill

This paper develops a framework to study the statistical power of revealed-preference tests. With randomly sampled budgets and mild smoothness of demand, statistical learning implies that any model consistent with the data must approximate…

Theoretical Economics · Economics 2026-02-12 Charles Gauthier , Raghav Malhotra , Agustin Troccoli Moretti

Performance prediction, the task of estimating a system's performance without performing experiments, allows us to reduce the experimental burden caused by the combinatorial explosion of different datasets, languages, tasks, and models. In…

Computation and Language · Computer Science 2021-02-11 Zihuiwen Ye , Pengfei Liu , Jinlan Fu , Graham Neubig

In multiple hypothesis testing, the volume of data, defined as the number of replications per null times the total number of nulls, usually defines the amount of resource required. On the other hand, power is an important measure of…

Statistics Theory · Mathematics 2009-06-05 Zhiyi Chi

An important issue for many economic experiments is how the experimenter can ensure sufficient power for rejecting one or more hypotheses. Here, we apply methods developed mainly within the area of clinical trials for testing multiple…

Methodology · Statistics 2021-08-06 Sebastian Jobjörnsson , Henning Schaak , Oliver Mußhoff , Tim Friede

We show that publishing results using the statistical significance filter---publishing only when the p-value is less than 0.05---leads to a vicious cycle of overoptimistic expectation of the replicability of results. First, we show…

Methodology · Statistics 2017-05-16 Shravan Vasishth , Andrew Gelman

How many experimental studies would have come to different conclusions had they been run on larger samples? I show how to estimate the expected number of statistically significant results that a set of experiments would have reported had…

Econometrics · Economics 2025-09-23 Stefan Faridani
‹ Prev 1 2 3 10 Next ›