English
Related papers

Related papers: False Variable Selection Rates in Regression

200 papers

The concept of $k$-FWER has received much attention lately as an appropriate error rate for multiple testing when one seeks to control at least $k$ false rejections, for some fixed $k\ge 1$. A less conservative notion, the $k$-FDR, has been…

Statistics Theory · Mathematics 2009-06-18 Sanat K. Sarkar , Wenge Guo

The highly influential two-group model in testing a large number of statistical hypotheses assumes that the test statistics are drawn independently from a mixture of a high probability null distribution and a low probability alternative.…

Methodology · Statistics 2020-12-08 Ruth Heller , Saharon Rosset

We propose a new method, semi-penalized inference with direct false discovery rate control (SPIDR), for variable selection and confidence interval construction in high-dimensional linear regression. SPIDR first uses a semi-penalized…

Methodology · Statistics 2013-12-02 Jian Huang , Shuangge Ma , Cun-Hui Zhang , Yong Zhou

Conventional feature selection algorithms applied to Pseudo Time-Series (PTS) data, which consists of observations arranged in sequential order without adhering to a conventional temporal dimension, often exhibit impractical computational…

Machine Learning · Computer Science 2024-03-14 Mohammad Rahman , Manzur Murshed , Shyh Wei Teng , Manoranjan Paul

High-dimensional variable selection, particularly in genomics, requires error-controlling procedures that scale to millions of predictors. The Terminating-Random Experiments (T-Rex) selector achieves false discovery rate (FDR) control by…

Methodology · Statistics 2026-04-10 Taulant Koka , Jasin Machkour , Daniel P. Palomar , Michael Muma

In many applications of multiple hypothesis testing where more than one false rejection can be tolerated, procedures controlling error rates measuring at least $k$ false rejections, instead of at least one, for some fixed $k\ge 1$ can…

Statistics Theory · Mathematics 2008-12-18 Sanat K. Sarkar

Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality requirement, while controlling the false discovery rate (FDR). Existing methods fix the…

Machine Learning · Computer Science 2026-04-20 Meiyi Zhu , Osvaldo Simeone

Multiple testing with false discovery rate (FDR) control has been widely conducted in the ``discrete paradigm" where p-values have discrete and heterogeneous null distributions. However, in this scenario existing FDR procedures often lose…

Methodology · Statistics 2019-07-23 Xiongzhi Chen , R. W. Doerge , Sanat K. Sarkar

This paper proposes a model-free and data-adaptive feature screening method for ultra-high dimensional datasets. The proposed method is based on the projection correlation which measures the dependence between two random vectors. This…

Methodology · Statistics 2021-02-16 Wanjun Liu , Yuan Ke , Jingyuan Liu , Runze Li

With the development of data collection techniques, analysis with a survival response and high-dimensional covariates has become routine. Here we consider an interaction model, which includes a set of low-dimensional covariates, a set of…

Methodology · Statistics 2023-11-27 Weijuan Liang , Qingzhao Zhang , Shuangge Ma

This paper motivates and develops a novel and focused approach to variable selection in linear regression models. For estimating the regression mean $\mu=\E\,(Y\midd x_0)$, for the covariate vector of a given individual, there is a list of…

Methodology · Statistics 2026-02-19 Nils Lid Hjort

Genomics biobanks are information treasure troves with thousands of phenotypes (e.g., diseases, traits) and millions of single nucleotide polymorphisms (SNPs). The development of methodologies that provide reproducible discoveries is…

Methodology · Statistics 2024-10-08 Jasin Machkour , Michael Muma , Daniel P. Palomar

The traditional approaches to false discovery rate (FDR) control in multiple hypothesis testing are usually based on the null distribution of a test statistic. However, all types of null distributions, including the theoretical,…

Methodology · Statistics 2021-04-13 Kun He , Mengjie Li , Yan Fu , Fuzhou Gong , Xiaoming Sun

This paper presents a survey on some recent advances for the type I error rate control in multiple testing methodology. We consider the problem of controlling the $k$-family-wise error rate (kFWER, probability to make $k$ false discoveries…

Methodology · Statistics 2011-03-15 Etienne Roquain

This paper develops a framework for testing for associations in a possibly high-dimensional linear model where the number of features/variables may far exceed the number of observational units. In this framework, the observations are split…

Methodology · Statistics 2018-05-04 Rina Foygel Barber , Emmanuel J. Candes

In many scenarios such as genome-wide association studies where dependences between variables commonly exist, it is often of interest to infer the interaction effects in the model. However, testing pairwise interactions among millions of…

Methodology · Statistics 2022-09-02 Jingyi Duan , Yang Ning , Xi Chen , Yong Chen

Obtaining high-quality labels for large datasets is expensive, requiring massive annotations from human experts. While AI models offer a cost-effective alternative by predicting labels, their label quality is compromised by the unavoidable…

Machine Learning · Computer Science 2026-02-17 Huipeng Huang , Wenbo Liao , Huajun Xi , Hao Zeng , Mengchen Zhao , Hongxin Wei

The False Discovery Rate (FDR) paradigm aims to attain certain control on Type I errors with relatively high power for multiple hypothesis testing. The Benjamini--Hochberg (BH) procedure is a well-known FDR controlling procedure. Under a…

Statistics Theory · Mathematics 2007-11-06 Zhiyi Chi

False discovery rate (FDR) is a cornerstone of modern multiple testing. However, it often fails to guarantee the reliability of "marginal" discoveries that lie at the boundary of the rejection set, which are often crucial in high-precision…

Methodology · Statistics 2026-05-12 Yifan Zhang , Wentao Zhang , Changliang Zou , Haojie Ren

The false discovery rate (FDR) and the false non-discovery rate (FNR), defined as the expected false discovery proportion (FDP) and the false non-discovery proportion (FNP), are the most popular benchmarks for multiple testing. Despite the…

Statistics Theory · Mathematics 2025-09-03 Yutong Nie , Yihong Wu