English
Related papers

Related papers: When More Is Less: Pitfalls of significance testin…

200 papers

We examine the role of trustworthiness and trust in statistical inference, arguing that it is the extent of trustworthiness in inferential statistical tools which enables trust in the conclusions. Certain tools, such as the p-value and…

Methodology · Statistics 2021-05-11 David J. Hand

There is a well-known problem in Null Hypothesis Significance Testing: many statistically significant results fail to replicate in subsequent experiments. We show that this problem arises because standard `point-form null' significance…

Methodology · Statistics 2025-02-06 Fintan Costello , Paul Watts

Statistical methods are indispensable to scientific inference. However, there exists a longstanding tension across a wide range of scientific disciplines about the role that ``context'' should play in the application of statistical methods…

Applications · Statistics 2026-04-06 Ashley I Naimi

P-values are widely used in both the social and natural sciences to quantify the statistical significance of observed results. The recent surge of big data research has made the p-value an even more popular tool to test the significance of…

Applications · Statistics 2023-01-05 Bertie Vidgen , Taha Yasseri

When researchers carry out a null hypothesis significance test, it is tempting to assume that a statistically significant result lowers Prob(H0), the probability of the null hypothesis being true. Technically, such a statement is…

Applications · Statistics 2022-04-19 Daniel J. Schad , Shravan Vasishth

Null hypothesis significance tests and p values are widely used despite very strong arguments against their use in many contexts. Confidence intervals are often recommended as an alternative, but these do not achieve the objective of…

Methodology · Statistics 2014-02-12 Michael Wood

A lot of Machine Learning (ML) and Deep Learning (DL) research is of an empirical nature. Nevertheless, statistical significance testing (SST) is still not widely used. This endangers true progress, as seeming improvements over a baseline…

Machine Learning · Computer Science 2022-04-15 Dennis Ulmer , Christian Hardmeier , Jes Frellsen

The logical and practical difficulties associated with research interpretation using P values and null hypothesis significance testing have been extensively documented. This paper describes an alternative, likelihood-based approach to…

Methodology · Statistics 2021-09-21 Nicholas Adams , Gerard O'Reilly

Statistical significance tests can provide evidence that the observed difference in performance between two methods is not due to chance. In Information Retrieval, some studies have examined the validity and suitability of such tests for…

Information Retrieval · Computer Science 2019-04-09 Javier Parapar , David E. Losada , Manuel A. Presedo-Quindimil , Alvaro Barreiro

As a convention, p-value is often computed in frequentist hypothesis testing and compared with the nominal significance level of 0.05 to determine whether or not to reject the null hypothesis. The smaller the p-value, the more significant…

Methodology · Statistics 2020-02-25 Haolun Shi , Guosheng Yin

It is quite common in modern research, for a researcher to test many hypotheses. The statistical (frequentist) hypothesis testing framework, does not scale with the number of hypotheses in the sense that naively performing many hypothesis…

Methodology · Statistics 2013-06-26 Jonathan Rosenblatt

The established language for statistical testing --- significance levels, power, and p-values --- is overly complicated and deceptively conclusive. Even teachers of statistics and scientists who use statistics misinterpret the results of…

Statistics Theory · Mathematics 2019-10-23 Glenn Shafer

In their recent comment, published in Nature, Jeffrey T.Leek and Roger D.Peng discuss how P-values are widely abused in null hypothesis significance testing . We agree completely with them and in this short comment we discuss the importance…

Quantum Physics · Physics 2015-05-26 Marian Kupczynski

After some general remarks about the interrelation between philosophical and statistical thinking, the discussion centres largely on significance tests. These are defined as the calculation of $p$-values rather than as formal procedures for…

Statistics Theory · Mathematics 2007-06-13 Deborah G. Mayo , D. R. Cox

Nurses should rely on the best evidence, but tend to struggle with statistics, impeding research integration into clinical practice. Statistical significance, a key concept in classical statistics, and its primary metric, the p-value, are…

Methodology · Statistics 2023-11-27 Christopher Holmberg

This paper offers a commentary on the use of notions of statistical significance in choice modelling. We review the reasons for uncertainty in parameter estimates, provide a precise discussion on the computation of measures of uncertainty…

Econometrics · Economics 2026-05-18 Stephane Hess , Andrew Daly , Michiel Bliemer , Angelo Guevara , Ricardo Daziano , Thijs Dekker

Statistical significance testing is used in natural language processing (NLP) to determine whether the results of a study or experiment are likely to be due to chance or if they reflect a genuine relationship. A key step in significance…

Computation and Language · Computer Science 2024-01-01 Palash Goyal , Qian Hu , Rahul Gupta

Null Hypothesis Statistical Testing is a dominant framework for conducting statistical analysis across the sciences. There remains considerable debate as to whether, and under what circumstances, evidence can be said to be confirmatory of a…

Statistics Theory · Mathematics 2024-05-28 Reid Dale

We show that publishing results using the statistical significance filter---publishing only when the p-value is less than 0.05---leads to a vicious cycle of overoptimistic expectation of the replicability of results. First, we show…

Methodology · Statistics 2017-05-16 Shravan Vasishth , Andrew Gelman

Experimental research on behavior and cognition frequently rests on stimulus or subject selection where not all characteristics can be fully controlled, even when attempting strict matching. For example, when contrasting patients to…

Methodology · Statistics 2016-08-29 Jona Sassenhagen , Phillip M. Alday