The Power of Tests for Detecting $p$-Hacking
Abstract
A flourishing empirical literature investigates the prevalence of -hacking based on the distribution of -values across studies. Interpreting results in this literature requires a careful understanding of the power of methods for detecting -hacking. We theoretically study the implications of likely forms of -hacking on the distribution of -values to understand the power of tests for detecting it. Power can be low and depends crucially on the -hacking strategy and the distribution of true effects. Combined tests for upper bounds and monotonicity and tests for continuity of the -curve tend to have the highest power for detecting -hacking.
Cite
@article{arxiv.2205.07950,
title = {The Power of Tests for Detecting $p$-Hacking},
author = {Graham Elliott and Nikolay Kudrin and Kaspar Wüthrich},
journal= {arXiv preprint arXiv:2205.07950},
year = {2025}
}
Comments
Some parts of this paper are based on material in earlier versions of our arXiv working paper "Detecting p-hacking" (arXiv:1906.06711), which were not included in the final published version (Elliott et al., 2022, Econometrica)