English
Related papers

Related papers: Gold after Randomized Sand: Model-X Split Knockoff…

200 papers

Simultaneously performing variable selection and inference in high-dimensional models is an open challenge in statistics and machine learning. The increasing availability of vast amounts of variables requires the adoption of specific…

Methodology · Statistics 2025-10-02 Marco Molinari , Magne Thoresen

We consider problems where many, somewhat redundant, hypotheses are tested and we are interested in reporting the most precise rejections, with false discovery rate (FDR) control. This is the case, for example, when researchers are…

Methodology · Statistics 2024-04-23 Paula Gablenz , Chiara Sabatti

A concept-based classifier can explain the decision process of a deep learning model by human-understandable concepts in image classification problems. However, sometimes concept-based explanations may cause false positives, which…

Machine Learning · Computer Science 2024-01-23 Kaiwen Xu , Kazuto Fukuchi , Youhei Akimoto , Jun Sakuma

In 2015, Barber and Candes introduced a new variable selection procedure called the knockoff filter to control the false discovery rate (FDR) and prove that this method achieves exact FDR control. Inspired by the work of Barber and Candes…

Methodology · Statistics 2019-07-23 Jiajie Chen , Anthony Hou , Thomas Y. Hou

We consider the problem of assessing the importance of multiple variables or factors from a dataset when side information is available. In principle, using side information can allow the statistician to pay attention to variables with a…

Methodology · Statistics 2020-01-23 Zhimei Ren , Emmanuel Candès

The knockoff filter of Barber and Candes (arXiv:1404.5609) is a flexible framework for multiple testing in supervised learning models, based on introducing synthetic predictor variables to control the false discovery rate (FDR). Using the…

Methodology · Statistics 2024-11-26 Yixiang Luo , William Fithian , Lihua Lei

Knockoffs is a new framework for controlling the false discovery rate (FDR) in multiple hypothesis testing problems involving complex statistical models. While there has been great emphasis on Type-I error control, Type-II errors have been…

Methodology · Statistics 2017-12-19 Asaf Weinstein , Rina Barber , Emmanuel Candes

Controlling the false discovery rate (FDR) is a powerful approach to multiple testing. In many applications, the tested hypotheses have an inherent hierarchical structure. In this paper, we focus on the fixed sequence structure where the…

Methodology · Statistics 2016-11-11 Gavin Lynch , Wenge Guo , Sanat K. Sarkar , Helmut Finner

Testing for differences in features between clusters in various applications often leads to inflated false positives when practitioners use the same dataset to identify clusters and then test features, an issue commonly known as ``double…

Methodology · Statistics 2024-10-10 Lijun Wang , Yingxin Lin , Hongyu Zhao

One challenge in exploratory association studies using observational data is that the associations between the predictors and the outcome are potentially weak and rare, and the candidate predictors have complex correlation structures. False…

Methodology · Statistics 2025-01-30 Runqiu Wang , Ran Dai , Hongying Dai , Evan French , Cheng Zheng

In high-dimensional data analysis, such as financial index tracking or biomedical applications, it is crucial to select the few relevant variables while maintaining control over the false discovery rate (FDR). In these applications, strong…

Portfolio Management · Quantitative Finance 2024-01-31 Jasin Machkour , Daniel P. Palomar , Michael Muma

The goal of feature selection is to identify important features that are relevant to explain an outcome variable. Most of the work in this domain has focused on identifying globally relevant features, which are features that are related to…

Machine Learning · Statistics 2019-05-30 Jaime Roquero Gimenez , James Zou

In this article, we propose a novel strategy for conducting variable selection without prior model topology knowledge using the knockoff method with boosted tree models. Our method is inspired by the original knockoff method, where the…

Methodology · Statistics 2020-02-24 Tao Jiang , Yuanyuan Li , Alison A. Motsinger-Reif

In high dimensional variable selection problems, statisticians often seek to design multiple testing procedures that control the False Discovery Rate (FDR), while concurrently identifying a greater number of relevant variables. Model-X…

Statistics Theory · Mathematics 2023-07-25 Taejoo Ahn , Licong Lin , Song Mei

While data-driven confounder selection requires careful consideration, it is frequently employed in observational studies. Widely recognized criteria for confounder selection include the minimal-set approach, which involves selecting…

Methodology · Statistics 2025-08-21 Kazuharu Harada , Masataka Taguri

Large-scale hypothesis testing is central to modern science, where controlling the False Discovery Rate (FDR) has become the standard approach to managing false positives across many simultaneous tests. Hypotheses rarely exist in isolation;…

Methodology · Statistics 2026-05-19 Binyamin Perets , Shie Mannor

False discovery rate (FDR) is a cornerstone of modern multiple testing. However, it often fails to guarantee the reliability of "marginal" discoveries that lie at the boundary of the rejection set, which are often crucial in high-precision…

Methodology · Statistics 2026-05-12 Yifan Zhang , Wentao Zhang , Changliang Zou , Haojie Ren

The robustness of fault detection algorithms against uncertainty is crucial in the real-world industrial environment. Recently, a new probabilistic design scheme called distributionally robust fault detection (DRFD) has emerged and received…

Optimization and Control · Mathematics 2026-01-16 Yulin Feng , Hailang Jin , Steven X. Ding , Hao Ye , Chao Shang

Conditional independence testing (CIT) is essential for reliable scientific discovery. It prevents spurious findings and enables controlled feature selection. Recent CIT methods have used machine learning (ML) models as surrogates of the…

Statistics Theory · Mathematics 2026-02-02 Angel Reyero-Lobo , Bertrand Thirion , Pierre Neuvial

We propose the use of a new false discovery rate (FDR) controlling procedure as a model selection penalized method, and compare its performance to that of other penalized methods over a wide range of realistic settings: nonorthogonal design…

Applications · Statistics 2009-05-19 Yoav Benjamini , Yulia Gavrilov