English
Related papers

Related papers: p-Hacking Inflates Type I Error Rates in the Error…

200 papers

Likelihood ratio tests are a widely used method in global analyses in particle physics. The computation of the statistical significance (p-value) of these tests is usually done with a simple formula that relies on Wilks' theorem. There are,…

High Energy Physics - Phenomenology · Physics 2013-08-21 Martin Wiebusch

One class of statistical hypothesis testing procedures is the indisputable equivalence tests, whose main objective is to establish practical equivalence rather than the usual statistical significant difference. These hypothesis tests are…

Methodology · Statistics 2024-01-04 Daniel Ochieng

In contrast to its common definition and calculation, interpretation of p-values diverges among statisticians. Since p-value is the basis of various methodologies, this divergence has led to a variety of test methodologies and evaluations…

Methodology · Statistics 2012-12-27 Tomokazu Konishi

The literature on hypothesis testing with data-dependent and post-hoc significance levels relies on a particular extension of the Type-I error to data-dependent levels. Existing arguments for this extension are heuristic, and primarily…

Statistics Theory · Mathematics 2026-05-28 Nick W. Koning

Inference is the process of using facts we know to learn about facts we do not know. A theory of inference gives assumptions necessary to get from the former to the latter, along with a definition for and summary of the resulting…

Machine Learning · Statistics 2021-09-27 Beau Coker , Cynthia Rudin , Gary King

While generative models, especially large language models (LLMs), are ubiquitous in today's world, principled mechanisms to assess their (in)correctness are limited. Using the conformal prediction framework, previous works construct sets of…

Machine Learning · Statistics 2026-04-02 Guneet S. Dhillon , Javier González , Teodora Pandeva , Alicia Curth

After some general remarks about the interrelation between philosophical and statistical thinking, the discussion centres largely on significance tests. These are defined as the calculation of $p$-values rather than as formal procedures for…

Statistics Theory · Mathematics 2007-06-13 Deborah G. Mayo , D. R. Cox

The validity of classical hypothesis testing requires the significance level $\alpha$ be fixed before any statistical analysis takes place. This is a stringent requirement. For instance, it prohibits updating $\alpha$ during (or after) an…

Statistics Theory · Mathematics 2026-01-21 Ben Chugg , Tyron Lardy , Aaditya Ramdas , Peter Grünwald

Fairness in machine learning (ML) is an ever-growing field of research due to the manifold potential for harm from algorithmic discrimination. To prevent such harm, a large body of literature develops new approaches to quantify fairness.…

Machine Learning · Computer Science 2024-01-10 Kristof Meding , Thilo Hagendorff

The literature on test set contamination largely focuses on detection, but the correction of contaminated test scores is underexplored. Our core proposal is to spike the training data by intentionally contaminating some test examples at…

Methodology · Statistics 2026-05-26 Johnny Tian-Zheng Wei , Jerry Li , Ameya Godbole , Robin Jia

We study a statistical framework for replicability based on a recently proposed quantitative measure of replication success, the sceptical $p$-value. A recalibration is proposed to obtain exact overall Type-I error control if the effect is…

Methodology · Statistics 2023-11-10 Charlotte Micheloud , Fadoua Balabdaoui , Leonhard Held

We examine the role of trustworthiness and trust in statistical inference, arguing that it is the extent of trustworthiness in inferential statistical tools which enables trust in the conclusions. Certain tools, such as the p-value and…

Methodology · Statistics 2021-05-11 David J. Hand

Increased availability of data and accessibility of computational tools in recent years have created unprecedented opportunities for scientific research driven by statistical analysis. Inherent limitations of statistics impose constrains on…

Genomics · Quantitative Biology 2016-09-13 Olga A. Vsevolozhskaya , Gabriel Ruiz , Dmitri V. Zaykin

Modern data analysis frequently involves large-scale hypothesis testing, which naturally gives rise to the problem of maintaining control of a suitable type I error rate, such as the false discovery rate (FDR). In many biomedical and…

Methodology · Statistics 2023-07-25 David S. Robertson , James M. S. Wason , Aaditya Ramdas

In high-dimensional linear models, the sparsity assumption is typically made, stating that most of the parameters are equal to zero. Under the sparsity assumption, estimation and, recently, inference have been well studied. However, in…

Methodology · Statistics 2019-07-09 Yinchu Zhu , Jelena Bradic

An important aspect of multiple hypothesis testing is controlling the significance level, or the level of Type I error. When the test statistics are not independent it can be particularly challenging to deal with this problem, without…

Statistics Theory · Mathematics 2009-03-04 Sandy Clarke , Peter Hall

This paper addresses the following general scenario: A scientist wishes to perform a battery of experiments, each generating a sequential stream of data, to investigate some phenomenon. The scientist would like to control the overall error…

Methodology · Statistics 2014-05-12 Jay Bartroff , Jinlin Song

In this paper, we investigate the impact of high-dimensional Principal Component (PC) adjustments on inferring the effects of variables on outcomes, with a focus on applications in genetic association studies where PC adjustment is commonly…

Statistics Theory · Mathematics 2025-06-30 Sohom Bhattacharya , Rounak Dey , Rajarshi Mukherjee

The classical theory for the meta-analysis of $p$-values is based on the assumption that if the overall null hypothesis is true, then all $p$-values used in a chosen combined test statistic are genuine, i.e., are observations from…

Computation · Statistics 2024-10-08 Rui Santos , M. Fátima Brilhante , Sandra Mendonça

While P-values are widely abused, they are a useful tool for many purposes; banning them is analogous to banning scalpels because most people do not know how to perform surgery. Many reported P-values are not genuine P-values, for a variety…

Other Statistics · Statistics 2022-04-19 Philip B. Stark