English
Related papers

Related papers: p-Hacking Inflates Type I Error Rates in the Error…

200 papers

In clinical studies upon which decisions are based there are two types of errors that can be made: a type I error arises when the decision is taken to declare a positive outcome when the truth is in fact negative, and a type II error arises…

Methodology · Statistics 2024-09-19 Andrew P Grieve

Statistical hypotheses are translations of scientific hypotheses into statements about one or more distributions, often concerning their centre. Tests that assess statistical hypotheses of centre implicitly assume a specific centre, e.g.,…

Methodology · Statistics 2024-02-21 Ryan Thompson , Catherine S. Forbes , Steven N. MacEachern , Mario Peruggia

A fundamental assumption of classical hypothesis testing is that the significance threshold $\alpha$ is chosen independently from the data. The validity of confidence intervals likewise relies on choosing $\alpha$ beforehand. We point out…

Applications · Statistics 2025-03-11 Jesse Hemerik , Nick W Koning

Hybrid clinical trials, that borrow real-world data (RWD), are gaining interest, especially for rare diseases. They assume RWD and randomized control arm be exchangeable, but violations can bias results, inflate type I error, or reduce…

A pervasive issue in statistical hypothesis testing is that the reported $p$-values are biased downward by data "peeking" -- the practice of reporting only progressively extreme values of the test statistic as more data samples are…

Statistics Theory · Mathematics 2020-11-04 Akshay Balsubramani

Principal Component Analysis (PCA) aims to find subspaces spanned by the so-called principal components that best represent the variance in the dataset. The deflation method is a popular meta-algorithm that sequentially finds individual…

Machine Learning · Computer Science 2024-05-30 Fangshuo Liao , Junhyung Lyle Kim , Cruz Barnum , Anastasios Kyrillidis

We show that adding noise before publishing data effectively screens $p$-hacked findings: spurious explanations produced by fitting many statistical models (data mining). Noise creates "baits" that affect two types of researchers…

Theoretical Economics · Economics 2024-05-21 Federico Echenique , Kevin He

A new formulation for the proportion of true null hypotheses $(\pi_0)$, based on the sum of all $p$-values and the average of expected $p$-value under the false null hypotheses has been proposed in the current work. This formulation of the…

Statistics Theory · Mathematics 2020-04-07 Aniket Biswas

Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…

Methodology · Statistics 2024-09-05 F. Richard Guo , Rajen D. Shah

This paper investigates type I error violations that occur when blinded sample size reviews are applied in equivalence testing. We give a derivation which explains why such violations are more pronounced in equivalence testing than in the…

Applications · Statistics 2021-09-08 Ekkehard Glimm , Lillian Yau , Heike Woehling

Background: Well-designed phase II trials must have acceptable error rates relative to a pre-specified success criterion, usually a statistically significant p-value. Such standard designs may not always suffice from a clinical perspective…

Applications · Statistics 2020-02-10 Satrajit Roychoudhury , Nicolas Scheuer , Beat Neuenschwander

Mathematics is a limited component of solutions to real-world problems, as it expresses only what is expected to be true if all our assumptions are correct, including implicit assumptions that are omnipresent and often incorrect.…

Methodology · Statistics 2023-09-14 Sander Greenland

Null Hypothesis Significance Testing is the \textit{de facto} tool for assessing effectiveness differences between Information Retrieval systems. Researchers use statistical tests to check whether those differences will generalise to online…

Information Retrieval · Computer Science 2025-07-23 David Otero , Javier Parapar , Álvaro Barreiro

We show that publishing results using the statistical significance filter---publishing only when the p-value is less than 0.05---leads to a vicious cycle of overoptimistic expectation of the replicability of results. First, we show…

Methodology · Statistics 2017-05-16 Shravan Vasishth , Andrew Gelman

Recently, it was shown that most popular IR measures are not interval-scaled, implying that decades of experimental IR research used potentially improper methods, which may have produced questionable results. However, it was unclear if and…

Information Retrieval · Computer Science 2021-01-08 Marco Ferrante , Nicola Ferro , Norbert Fuhr

In a recent simulation study, Goodman et al. (2019) compare several methods with regard to their type I and type II error rates in case of a thick null hypothesis that includes all values that are practically equivalent to the point null…

Methodology · Statistics 2022-06-07 Robin Tim Dreher , Leona Hoffmann , Arne Kramer-Sunderbrink , Peter Pütz , Robin Werner

This paper studies the construction of p-values for nonparametric outlier detection, taking a multiple-testing perspective. The goal is to test whether new independent samples belong to the same distribution as a reference data set or are…

Methodology · Statistics 2024-03-12 Stephen Bates , Emmanuel Candès , Lihua Lei , Yaniv Romano , Matteo Sesia

We consider the problem of testing for differences in group-specific slopes between the selected groups in panel data identified via k-means clustering. In this setting, the classical Wald-type test statistic is problematic because it…

Methodology · Statistics 2025-11-07 Chuang Wan , Jiajun Sun , Xingbai Xu

How much does a machine learning algorithm leak about its training data, and why? Membership inference attacks are used as an auditing tool to quantify this leakage. In this paper, we present a comprehensive \textit{hypothesis testing…

Machine Learning · Computer Science 2022-09-14 Jiayuan Ye , Aadyaa Maddi , Sasi Kumar Murakonda , Vincent Bindschaedler , Reza Shokri

Since its debut in the 18th century, the P-value has been an important part of hypothesis testing-based scientific discoveries. As the statistical engine accelerates, questions are beginning to be raised, asking to what extent scientific…

‹ Prev 1 3 4 5 6 7 10 Next ›