English
Related papers

Related papers: P-values: misunderstood and misused

200 papers

The role of statisticians in society is to provide tools, techniques, and guidance with regards to how much to trust data. This role is increasingly more important with more data and more misinformation than ever before. The American…

Physics and Society · Physics 2020-07-08 Joshua T. Vogelstein

When a scientist performs an experiment they normally acquire a set of measurements and are expected to demonstrate that their results are "statistically significant" thus confirming whatever hypothesis they are testing. The main method for…

Other Statistics · Statistics 2011-09-30 Jacob Levman

This paper re-visits the problem of deciding between two simple hypotheses, the setting considered by Neyman and Pearson in developing their fundamental lemma. It studies the decision process induced by the most powerful test and the…

Statistics Theory · Mathematics 2019-11-19 Edsel A. Pena

We discuss problems the null hypothesis significance testing (NHST) paradigm poses for replication and more broadly in the biomedical and social sciences as well as how these problems remain unresolved by proposals involving modified…

Methodology · Statistics 2021-07-21 Blakeley B. McShane , David Gal , Andrew Gelman , Christian Robert , Jennifer L. Tackett

It is quite common in modern research, for a researcher to test many hypotheses. The statistical (frequentist) hypothesis testing framework, does not scale with the number of hypotheses in the sense that naively performing many hypothesis…

Methodology · Statistics 2013-06-26 Jonathan Rosenblatt

The mid-p-value is a proposed improvement on the ordinary p-value for the case where the test statistic is partially or completely discrete. In this case, the ordinary p-value is conservative, meaning that its null distribution is larger…

Statistics Theory · Mathematics 2017-06-02 Patrick Rubin-Delanchy , Nicholas A. Heard , Daniel John Lawson

In this study, we propose a two-stage procedure for hypothesis testing, where the first stage is conventional hypothesis testing and the second is an equivalence testing procedure using an introduced Empirical Equivalence Bound. In 2016,…

Methodology · Statistics 2020-01-01 Yi Zhao , Brian S. Caffo , Joshua B. Ewen

We are concerned with testing replicability hypotheses for many endpoints simultaneously. This constitutes a multiple test problem with composite null hypotheses. Traditional $p$-values, which are computed under least favourable parameter…

Methodology · Statistics 2020-02-26 Anh-Tuan Hoang , Thorsten Dickhaus

In this paper, we demonstrate that a new measure of evidence we developed called the Dempster-Shafer p-value which allow for insights and interpretations which retain most of the structure of the p-value while covering for some of the…

Methodology · Statistics 2024-02-28 Kentaro Hoffman , Kai Zhang , Tyler McCormick , Jan Hannig

This article develops $p$-values for evaluating means of normal populations that make use of indirect or prior information. A $p$-value of this type is based on a biased test statistic that is optimal on average with respect to a…

Methodology · Statistics 2019-12-12 Peter D. Hoff

The higher criticism of a family of tests starts with the individual uncorrected p-values of each test. It then requires a procedure for deciding whether the collection of p-values indicates the presence of a real effect and if possible…

Statistical significance testing of differences in values of metrics like recall, precision and balanced F-score is a necessary part of empirical natural language processing. Unfortunately, we find in a set of experiments that many commonly…

Computation and Language · Computer Science 2007-05-23 Alexander Yeh

There is recent interest in estimating the false discovery rate (FDR) with published p-values. However, there is little formal research that addresses the manner and extent to which the presumed selection, or publication, bias model impacts…

Methodology · Statistics 2026-03-03 Tianyu Cao , Sangyoon Yi , Joshua Habiger

We seek to conduct statistical inference for a large collection of primary parameters, each with its own nuisance parameters. Our approach is partially Bayesian, in that we treat the primary parameters as fixed while we model the nuisance…

Methodology · Statistics 2025-12-10 Nikolaos Ignatiadis , Li Ma

Quantifying uncertainty in detected changepoints is an important problem. However it is challenging as the naive approach would use the data twice, first to detect the changes, and then to test them. This will bias the test, and can lead to…

Methodology · Statistics 2026-05-11 Rachel Carrington , Paul Fearnhead

When testing multiple hypothesis in a survey --e.g. many different source locations, template waveforms, and so on-- the final result consists in a set of confidence intervals, each one at a desired confidence level. But the probability…

General Relativity and Quantum Cosmology · Physics 2009-11-11 L. Baggio , G. A. Prodi

Hypothesis testing is a central statistical method in psychological research and the cognitive sciences. While the problems of null hypothesis significance testing (NHST) have been debated widely, few attractive alternatives exist. In this…

Methodology · Statistics 2020-06-08 Riko Kelter , Julio Michael Stern

Importance sampling is a common technique for Monte Carlo approximation, including Monte Carlo approximation of p-values. Here it is shown that a simple correction of the usual importance sampling p-values creates valid p-values, meaning…

Computation · Statistics 2011-04-12 Matthew T. Harrison

The American Statistical Association (ASA) statement on statistical significance and P-values \cite{wasserstein2016asa} cautioned statisticians against making scientific decisions solely on the basis of traditional P-values. The statement…

Methodology · Statistics 2024-02-22 Abhisek Chakraborty , Megan H. Murray , Ilya Lipkovich , Yu Du

The energy test method is a multi-dimensional test of whether two samples are consistent with arising from the same underlying population, through the calculation of a single test statistic (called the $T$-value). The method has recently…

Data Analysis, Statistics and Probability · Physics 2018-04-19 W. Barter , C. Burr , C. Parkes