English
Related papers

Related papers: Improving the adjusted Benjamini--Hochberg method …

200 papers

Estimation of heterogeneous treatment effects is an essential component of precision medicine. Model and algorithm-based methods have been developed within the causal inference framework to achieve valid estimation and inference. Existing…

Methodology · Statistics 2021-05-10 Ruohong Li , Honglang Wang , Wanzhu Tu

Many modern applications require using data to select the statistical tasks and make valid inference after selection. In this article, we provide a unifying approach to control for a class of selective risks. Our method is motivated by a…

Methodology · Statistics 2024-11-11 Zijun Gao , Wenjie Hu , Qingyuan Zhao

Traditional variable selection methods could fail to be sign consistent when irrepresentable conditions are violated. This is especially critical in high-dimensional settings when the number of predictors exceeds the sample size. In this…

Methodology · Statistics 2022-04-26 Fei Xue , Annie Qu

Controlled variable selection is an important analytical step in various scientific fields, such as brain imaging or genomics. In these high-dimensional data settings, considering too many variables leads to poor models and high costs,…

Methodology · Statistics 2023-10-17 Alexandre Blain , Bertrand Thirion , Olivier Grisel , Pierre Neuvial

False discovery rate (FDR) controlling procedures provide important statistical guarantees for the replicability in signal identification based on multiple hypotheses testing. In many fields of study, FDR controlling procedures are used in…

Methodology · Statistics 2022-10-04 Ran Dai , Cheng Zheng

In many research fields, researchers aim to identify significant associations between a set of explanatory variables and a response while controlling the FDR. The Knockoff filter has been recently proposed in the frequentist paradigm to…

Methodology · Statistics 2026-04-22 Lorenzo Focardi-Olmi , Anna Gottard , Michele Guindani , Marina Vannucci

We propose a new empirical Bayes method for covariate-assisted multiple testing with false discovery rate (FDR) control, where we model the local false discovery rate for each hypothesis as a function of both its covariates and p-value. Our…

Methodology · Statistics 2021-07-01 Patrick Chao , William Fithian

The Model-X knockoff procedure has recently emerged as a powerful approach for feature selection with statistical guarantees. The advantage of knockoff is that if we have a good model of the features X, then we can identify salient features…

Machine Learning · Statistics 2019-05-30 Jaime Roquero Gimenez , James Zou

Hybrid controlled trials (HCTs), which augment randomized controlled trials (RCTs) with external controls (ECs), are increasingly receiving attention as a way to address limited power, slow accrual, and ethical concerns in clinical…

Methodology · Statistics 2025-05-02 Jiajun Liu , Ke Zhu , Shu Yang , Xiaofei Wang

Preserving model fidelity is essential for stealthy text-to-image (T2I) backdoor attacks. Existing methods such as Learning without Forgetting (LwF) rely on output-based distillation, which provides limited regularization. We introduce…

Machine Learning · Computer Science 2026-05-12 Lu Bowen , Xinyu Tang , Yin Yin Low , Shu-Min Leong

Although there is a huge literature on feature selection for the Cox model, none of the existing approaches can control the false discovery rate (FDR) unless the sample size tends to infinity. In addition, there is no formal power analysis…

Methodology · Statistics 2023-08-02 Daoji Li , Jinzhao Yu , Hui Zhao

Differential privacy provides a rigorous framework for privacy-preserving data analysis. This paper proposes the first differentially private procedure for controlling the false discovery rate (FDR) in multiple hypothesis testing. Inspired…

Statistics Theory · Mathematics 2021-07-06 Cynthia Dwork , Weijie J. Su , Li Zhang

A new online multiple testing procedure is described in the context of anomaly detection, which controls the False Discovery Rate (FDR). An accurate anomaly detector must control the false positive rate at a prescribed level while keeping…

Methodology · Statistics 2024-12-17 Etienne Krönert , Alain Célisse , Dalila Hattab

The rapid generation of complex, highly skewed, and zero-inflated multi-source count data poses significant challenges for variable selection, particularly in biomedical domains like tumor development and metabolic dysregulation. To address…

Applications · Statistics 2025-11-11 Shan Tang , Shanjun Mao , Shourong Ma , Falong Tan

We address challenges in variable selection with highly correlated data that are frequently present in finance, economics, but also in complex natural systems as e.g. weather. We develop a robustified version of the knockoff framework,…

Econometrics · Economics 2022-06-14 Konstantin Görgen , Abdolreza Nazemi , Melanie Schienle

Clinical trials with a hybrid control arm (a control arm constructed from a combination of randomized patients and real-world data on patients receiving usual care in standard clinical practice) have the potential to decrease the cost of…

Methodology · Statistics 2021-08-20 Joanna Harton , Brian Segal , Ronac Mamtani , Nandita Mitra , Rebecca Hubbard

Many methods have been developed to estimate the set of relevant variables in a sparse linear model Y= XB+e where the dimension p of B can be much higher than the length n of Y. Here we propose two new methods based on multiple hypotheses…

Statistics Theory · Mathematics 2012-06-12 Florian Rohart

Motivated by the CATHGEN data, we develop a new statistical learning method for simultaneous variable selection and parameter estimation under the context of generalized partly linear models for data with high-dimensional covariates. The…

Methodology · Statistics 2023-11-02 Christian Chan , Xiaotian Dai , Thierry Chekouo , Quan Long , Xuewen Lu

Obtaining high-quality labels for large datasets is expensive, requiring massive annotations from human experts. While AI models offer a cost-effective alternative by predicting labels, their label quality is compromised by the unavoidable…

Machine Learning · Computer Science 2026-02-17 Huipeng Huang , Wenbo Liao , Huajun Xi , Hao Zeng , Mengchen Zhao , Hongxin Wei

Ordered testing procedures are multiple testing procedures that exploit a pre-specified ordering of the null hypotheses, from most to least promising. We analyze and compare the power of several recent proposals using the asymptotic…

Methodology · Statistics 2016-06-08 Lihua Lei , William Fithian