English
Related papers

Related papers: Rate optimal multiple testing procedure in high-di…

200 papers

We consider selecting the top-$m$ alternatives from a finite number of alternatives via Monte Carlo simulation. Under a Bayesian framework, we formulate the sampling decision as a stochastic dynamic programming problem, and develop a…

Optimization and Control · Mathematics 2023-08-22 Gongbo Zhang , Yijie Peng , Jianghua Zhang , Enlu Zhou

Assuming that data are collected sequentially from independent streams, we consider the simultaneous testing of multiple binary hypotheses under two general setups; when the number of signals (correct alternatives) is known in advance, and…

Statistics Theory · Mathematics 2017-02-14 Yanglei Song , Georgios Fellouris

We attempt to recover an $n$-dimensional vector observed in white noise, where $n$ is large and the vector is known to be sparse, but the degree of sparsity is unknown. We consider three different ways of defining sparsity of a vector:…

Statistics Theory · Mathematics 2007-06-13 Felix Abramovich , Yoav Benjamini , David L. Donoho , Iain M. Johnstone

In a high dimensional regression setting in which the number of variables ($p$) is much larger than the sample size ($n$), the number of possible two-way interactions between the variables is immense. If the number of variables is in the…

Methodology · Statistics 2024-06-26 Marianne A Jonker , Luc van Schijndel , Eric Cator

Multiple tests are designed to test a whole collection of null hypotheses simultaneously. Their quality is often judged by the false discovery rate (FDR), i.e. the expectation of the quotient of the number of false rejections divided by the…

Statistics Theory · Mathematics 2015-11-24 Julia Benditkis , Philipp Heesen , Arnold Janssen

We propose sufficient conditions and computationally efficient procedures for false discovery rate control in multiple testing when the $p$-values are related by a known \emph{dependency graph} -- meaning that we assume independence of…

Methodology · Statistics 2025-07-01 Drew T. Nguyen , William Fithian

This paper addresses the challenge of efficiently capturing a high proportion of true signals for subsequent data analyses when sample sizes are relatively limited with respect to data dimension. We propose the signal missing rate as a new…

Methodology · Statistics 2018-08-30 X. Jessie Jeng , Teng Zhang , Jung-Ying Tzeng

In this paper, we introduce an innovative testing procedure for assessing individual hypotheses in high-dimensional linear regression models with measurement errors. This method remains robust even when either the X-model or Y-model is…

Methodology · Statistics 2025-01-14 Shijie Cui , Xu Guo , Songshan Yang , Zhe Zhang

A topological multiple testing scheme for one-dimensional domains is proposed where, rather than testing every spatial or temporal location for the presence of a signal, tests are performed only at the local maxima of the smoothed observed…

Statistics Theory · Mathematics 2012-03-15 Armin Schwartzman , Yulia Gavrilov , Robert J. Adler

This research deals with massive multiple hypothesis testing. First regarding multiple tests as an estimation problem under a proper population model, an error measurement called Erroneous Rejection Ratio (ERR) is introduced and related to…

Statistics Theory · Mathematics 2007-06-13 Cheng Cheng

In many scenarios such as genome-wide association studies where dependences between variables commonly exist, it is often of interest to infer the interaction effects in the model. However, testing pairwise interactions among millions of…

Methodology · Statistics 2022-09-02 Jingyi Duan , Yang Ning , Xi Chen , Yong Chen

The genetic basis of multiple phenotypes such as gene expression, metabolite levels, or imaging features is often investigated by testing a large collection of hypotheses, probing the existence of association between each of the traits and…

Applications · Statistics 2015-04-06 Christine Peterson , Marina Bogomolov , Yoav Benjamini , Chiara Sabatti

Controlling false discovery rate (FDR) while leveraging the side information of multiple hypothesis testing is an emerging research topic in modern data science. Existing methods rely on the test-level covariates while ignoring possible…

Machine Learning · Statistics 2021-01-26 Lin Qiu , Nils Murrugarra-Llerena , Vítor Silva , Lin Lin , Vernon M. Chinchilli

Out of the participants in a randomized experiment with anticipated heterogeneous treatment effects, is it possible to identify which subjects have a positive treatment effect? While subgroup analysis has received attention, claims about…

Methodology · Statistics 2024-05-14 Boyan Duan , Larry Wasserman , Aaditya Ramdas

The paper considers variable selection in linear regression models where the number of covariates is possibly much larger than the number of observations. High dimensionality of the data brings in many complications, such as (possibly…

Methodology · Statistics 2016-11-29 Haeran Cho , Piotr Fryzlewicz

False discovery rate (FDR) procedures provide misleading inference when testing multiple null hypotheses with heterogeneous multinomial data. For example, in the motivating study the goal is to identify species of bacteria near the roots of…

Methodology · Statistics 2015-11-05 Joshua Habiger , David Watts , Michael Anderson

The false discovery rate (FDR) and false nondiscovery rate (FNDR) have received considerable attention in the literature on multiple testing. These performance measures are also appropriate for classification, and in this work we develop…

Statistics Theory · Mathematics 2009-01-28 Clayton Scott , Gowtham Bellala , Rebecca Willett

We investigate the performance of a family of multiple comparison procedures for strong control of the False Discovery Rate ($\mathsf{FDR}$). The $\mathsf{FDR}$ is the expected False Discovery Proportion ($\mathsf{FDP}$), that is, the…

Statistics Theory · Mathematics 2008-11-21 Pierre Neuvial

This paper considers multiple regression procedures for analyzing the relationship between a response variable and a vector of covariates in a nonparametric setting where both tuning parameters and the number of covariates need to be…

Statistics Theory · Mathematics 2007-06-13 Chad M. Schafer , Kjell A. Doksum

High-dimensional, low sample-size (HDLSS) data problems have been a topic of immense importance for the last couple of decades. There is a vast literature that proposed a wide variety of approaches to deal with this situation, among which…

Methodology · Statistics 2021-07-09 Kaixu Yang , Tapabrata Maiti