English
Related papers

Related papers: False Discovery Rate Control via Data Splitting fo…

200 papers

Modern biomedical research frequently involves testing multiple related hypotheses, while maintaining control over a suitable error rate. In many applications the false discovery rate (FDR), which is the expected proportion of false…

Methodology · Statistics 2018-09-27 David S. Robertson , James M. S. Wason

Advanced computing and data acquisition technologies have made possible the collection of high-dimensional data streams in many fields. Efficient online monitoring tools which can correctly identify any abnormal data stream for such data…

Methodology · Statistics 2017-12-15 Jun Li

Multiple tests are designed to test a whole collection of null hypotheses simultaneously. Their quality is often judged by the false discovery rate (FDR), i.e. the expectation of the quotient of the number of false rejections divided by the…

Statistics Theory · Mathematics 2015-11-24 Julia Benditkis , Philipp Heesen , Arnold Janssen

With the rapid growth of crowdsourcing platforms it has become easy and relatively inexpensive to collect a dataset labeled by multiple annotators in a short time. However due to the lack of control over the quality of the annotators, some…

Machine Learning · Statistics 2016-06-17 Qianqian Xu , Jiechao Xiong , Xiaochun Cao , Yuan Yao

This paper presents a powerful methodology for flexible full-data nonparametric novelty detection that offers distribution-free false discovery rate (FDR) control guarantees. Building on the full conformal inference framework and the…

Methodology · Statistics 2026-04-21 Junu Lee , Ilia Popov , Zhimei Ren

Genomics biobanks are information treasure troves with thousands of phenotypes (e.g., diseases, traits) and millions of single nucleotide polymorphisms (SNPs). The development of methodologies that provide reproducible discoveries is…

Methodology · Statistics 2024-10-08 Jasin Machkour , Michael Muma , Daniel P. Palomar

The introduction of the false discovery rate (FDR) by Benjamini and Hochberg has spurred a great interest in developing methodologies to control the FDR in various settings. The majority of existing approaches, however, address the FDR…

Methodology · Statistics 2016-06-09 Kasra Alishahi , Ahmad Reza Ehyaei , Ali Shojaie

The False Discovery Rate (FDR) paradigm aims to attain certain control on Type I errors with relatively high power for multiple hypothesis testing. The Benjamini--Hochberg (BH) procedure is a well-known FDR controlling procedure. Under a…

Statistics Theory · Mathematics 2007-11-06 Zhiyi Chi

Inequalities are key tools to prove FDR control of a multiple test. The present paper studies upper and lower bounds for the FDR under various dependence structures of p-values, namely independence, reverse martingale dependence and…

Statistics Theory · Mathematics 2015-02-18 Philipp Heesen , Arnold Janssen

In many large scale multiple testing applications, the hypotheses often have a known graphical structure, such as gene ontology in gene expression data. Exploiting this graphical structure in multiple testing procedures can improve power as…

Methodology · Statistics 2018-12-04 Wenge Guo , Gavin Lynch , Joseph P. Romano

Identifying signals that replicate across multiple studies is essential for establishing robust scientific evidence, yet existing methods for high-dimensional replicability analysis either rely on restrictive modeling assumptions, are…

Methodology · Statistics 2026-03-05 Haochen Lei , Yan Li , Hongyuan Cao

Multiple comparison procedures that control a family-wise error rate or false discovery rate provide an achieved error rate as the adjusted p-value for each hypothesis tested. However, since such p-values are not probabilities that the null…

Methodology · Statistics 2013-09-03 David R. Bickel

We present a novel necessary and sufficient principle for multiple testing methods controlling an expected loss. This principle asserts that every such multiple testing method is a special case of a general closed testing procedure based on…

Methodology · Statistics 2026-01-05 Ziyu Xu , Aldo Solari , Lasse Fischer , Rianne de Heide , Aaditya Ramdas , Jelle Goeman

Modern scientific technology has provided a new class of large-scale simultaneous inference problems, with thousands of hypothesis tests to consider at the same time. Microarrays epitomize this type of technology, but similar situations…

Statistics Theory · Mathematics 2007-11-06 Bradley Efron

This paper studies the distributed conditional feature screening for massive data with ultrahigh-dimensional features. Specifically, three distributed partial correlation feature screening methods (SAPS, ACPS and JDPS methods) are firstly…

Methodology · Statistics 2024-03-12 Naiwen Pang , Xiaochao Xia

Post-clustering inference in single-cell RNA sequencing (scRNA-seq) analysis presents significant challenges in controlling Type I error during differential expression analysis. Data fission, a promising approach that aims to split data…

Methodology · Statistics 2026-03-23 Benjamin Hivert , Denis Agniel , Rodolphe Thiébaut , Boris P. Hejblum

This paper explores the multiple testing problem for sparse high-dimensional data with binary outcomes. We propose novel empirical Bayes multiple testing procedures based on a spike-and-slab posterior and then evaluate their performance in…

Statistics Theory · Mathematics 2025-06-16 Yu-Chien Bo Ning

We attempt to recover an $n$-dimensional vector observed in white noise, where $n$ is large and the vector is known to be sparse, but the degree of sparsity is unknown. We consider three different ways of defining sparsity of a vector:…

Statistics Theory · Mathematics 2007-06-13 Felix Abramovich , Yoav Benjamini , David L. Donoho , Iain M. Johnstone

This paper continues the line of research initiated in Liu et. al. (2016) on developing a novel framework for multiple testing of hypotheses grouped in a one-way classified form using hypothesis-specific local false discovery rates…

Methodology · Statistics 2022-11-22 Sanat K. Sarkar , Zhigen Zhao

Large-scale multiple two-sample {\em Student}'s $t$ testing problems often arise from the statistical analysis of scientific data. To detect components with different values between two mean vectors, a well-known procedure is to apply the…

Methodology · Statistics 2014-10-17 Weidong Liu