English
Related papers

Related papers: When to adjust alpha during multiple testing: A co…

200 papers

As large language models (LLMs) have grown in prevalence, particular benchmarks have become essential for the evaluation of these models and for understanding model capabilities. Most commonly, we use test accuracy averaged across multiple…

Computation and Language · Computer Science 2024-11-12 Vipul Gupta , David Pantoja , Candace Ross , Adina Williams , Megan Ung

Valid statistical inference is challenging when the sample is subject to unknown selection bias. Data integration can be used to correct for selection bias when we have a parallel probability sample from the same population with some common…

Methodology · Statistics 2023-07-24 Zhonglei Wang , Shu Yang , Jae Kwang Kim

Hypothesis testing is one of the most common types of data analysis and forms the backbone of scientific research in many disciplines. Analysis of variance (ANOVA) in particular is used to detect dependence between a categorical and a…

Cryptography and Security · Computer Science 2019-03-05 Marika Swanberg , Ira Globus-Harris , Iris Griffith , Anna Ritz , Adam Groce , Andrew Bray

In this paper we discuss testing for an interaction in the two-way ANOVA with just one observation per cell. The known results are reviewed and a simulation study is performed to evaluate type I and type II risks of the tests. It is shown…

Statistics Theory · Mathematics 2012-07-13 Petr Simecek , Marie Simeckova

Hypothesis tests under order restrictions arise in a wide range of scientific applications. By exploiting inequality constraints, such tests can achieve substantial gains in power and interpretability. However, these gains come at a cost:…

Methodology · Statistics 2026-02-18 Ori Davidov

Bonferroni's correction is a popular tool to address multiplicity but is notorious for its low power when tests are dependent. This paper proposes a practical modification of Bonferroni's correction when test statistics are jointly normal…

Methodology · Statistics 2026-02-24 Caleb Hiltunen , Yeonwoo Rho

Incompatible, i.e. non-jointly measurable quantum measurements are a necessary resource for many information processing tasks. It is known that increasing the number of distinct measurements usually enhances the incompatibility of a…

Quantum Physics · Physics 2023-09-28 Lucas Tendick , Hermann Kampermann , Dagmar Bruß

This paper raises concerns about the advantages of using statistical significance tests in research assessments as has recently been suggested in the debate about proper normalization procedures for citation indicators. Statistical…

Digital Libraries · Computer Science 2012-09-26 Jesper W. Schneider

Data with multiple functional recordings at each observational unit are increasingly common in various fields including medical imaging and environmental sciences. To conduct inference for such observations, we develop a paired two-sample…

Methodology · Statistics 2025-06-16 Colin Decker , Dehan Kong , Stanislav Volgushev

Imbalances in covariates between treatment groups are frequent in observational studies and can lead to biased comparisons. Various adjustment methods can be employed to correct these biases in the context of multi-level treatments ($>$ 2).…

Applications · Statistics 2021-06-04 Diop S. Arona , Duchesne Thierry , Cumming Steven , Diop Awa , Talbot Denis

Recently, it was shown that most popular IR measures are not interval-scaled, implying that decades of experimental IR research used potentially improper methods, which may have produced questionable results. However, it was unclear if and…

Information Retrieval · Computer Science 2021-01-08 Marco Ferrante , Nicola Ferro , Norbert Fuhr

The field of psychological sciences has been grappling with the replicability crisis. Various issues have been identified as potential sources of this problem. We bring to light a potential source that has largely been overlooked and…

Methodology · Statistics 2025-04-28 Yoav Zeevi , Sofi Astashenko , Liad Mudrik , Yoav Benjamini

In the evaluation of treatment effects, it is of major policy interest to know if the treatment is beneficial for some and harmful for others, a phenomenon known as qualitative interaction. We formulate this question as a multiple testing…

Methodology · Statistics 2017-08-29 Qingyuan Zhao , Dylan S. Small , Weijie Su

This paper investigates type I error violations that occur when blinded sample size reviews are applied in equivalence testing. We give a derivation which explains why such violations are more pronounced in equivalence testing than in the…

Applications · Statistics 2021-09-08 Ekkehard Glimm , Lillian Yau , Heike Woehling

We adapt Higher Criticism (HC) to the comparison of two frequency tables which may -- or may not -- exhibit moderate differences between the tables in some unknown, relatively small subset out of a large number of categories. Our analysis…

Statistics Theory · Mathematics 2023-08-29 David L. Donoho , Alon Kipnis

Modern statisticians are often presented with hundreds or thousands of hypothesis testing problems to evaluate at the same time, generated from new scientific technologies such as microarrays, medical and satellite imaging devices, or flow…

Applications · Statistics 2008-12-18 Bradley Efron

In this paper we introduce a novel procedure for improving multiple testing procedures (MTPs) under scenarios when the null hypothesis $p$-values tend to be stochastically larger than standard uniform (referred to as 'inflated'). An…

Methodology · Statistics 2025-08-29 Jules L. Ellis , Jakub Pecanka , Jelle Goeman

Hypothesis testing is a key part of empirical science and multiple testing as well as the combination of evidence from several tests are continued areas of research. In this article we consider the problem of combining the results of…

Statistics Theory · Mathematics 2022-07-15 Phillip B. Mogensen , Bo Markussen

Statistical tests that compare classification algorithms are univariate and use a single performance measure, e.g., misclassification error, $F$ measure, AUC, and so on. In multivariate tests, comparison is done using multiple measures…

Machine Learning · Statistics 2014-09-17 Olcay Taner Yildiz , Ethem Alpaydin

We consider a permutation method for testing whether observations given in their natural pairing exhibit an unusual level of similarity in situations where any two observations may be similar at some unknown baseline level. Under a null…

Statistics Theory · Mathematics 2007-06-13 Larry Goldstein , Yosef Rinott