English
Related papers

Related papers: CoxKnockoff: Controlled Feature Selection for the …

200 papers

We present a novel method for controlling the $k$-familywise error rate ($k$-FWER) in the linear regression setting using the knockoffs framework first introduced by Barber and Cand\`es. Our procedure, which we also refer to as knockoffs,…

Methodology · Statistics 2015-11-10 Lucas Janson , Weijie Su

False discovery rate (FDR) control is a popular approach for maintaining the integrity of statistical analyses, especially in high-dimensional data settings, where multiple comparisons increase the risk of false positives. FDR control has…

Signal Processing · Electrical Eng. & Systems 2026-03-03 Fabian Scheidt , Jasin Machkour , Michael Muma

Simultaneously performing variable selection and inference in high-dimensional regression models is an open challenge in statistics and machine learning. The increasing availability of vast amounts of variables requires the adoption of…

Methodology · Statistics 2025-05-08 Marco Molinari , Magne Thoresen

We tackle the problem of selecting from among a large number of variables those that are 'important' for an outcome. We consider situations where groups of variables are also of interest in their own right. For example, each variable might…

Methodology · Statistics 2018-08-13 Eugene Katsevich , Chiara Sabatti

Competition-based approach to controlling the false discovery rate (FDR) recently rose to prominence when, generalizing it to sequential hypothesis testing, Barber and Cand\`es used it as part of their knockoff-filter. Control of the FDR…

Methodology · Statistics 2023-02-24 Arya Ebadi , Dong Luo , Jack Freestone , William Stafford Noble , Uri Keich

We consider problems where many, somewhat redundant, hypotheses are tested and we are interested in reporting the most precise rejections, with false discovery rate (FDR) control. This is the case, for example, when researchers are…

Methodology · Statistics 2024-04-23 Paula Gablenz , Chiara Sabatti

We propose the use of a new false discovery rate (FDR) controlling procedure as a model selection penalized method, and compare its performance to that of other penalized methods over a wide range of realistic settings: nonorthogonal design…

Applications · Statistics 2009-05-19 Yoav Benjamini , Yulia Gavrilov

We introduce DiffKnock, a diffusion-based knockoff framework for high-dimensional feature selection with finite-sample false discovery rate (FDR) control. DiffKnock addresses two key limitations of existing knockoff methods: preserving…

Methodology · Statistics 2025-10-03 Heng Ge , Qing Lu

In high dimensional variable selection problems, statisticians often seek to design multiple testing procedures that control the False Discovery Rate (FDR), while concurrently identifying a greater number of relevant variables. Model-X…

Statistics Theory · Mathematics 2023-07-25 Taejoo Ahn , Licong Lin , Song Mei

Model-X knockoffs allows analysts to perform feature selection using almost any machine learning algorithm while still provably controlling the expected proportion of false discoveries. To apply model-X knockoffs, one must construct…

Methodology · Statistics 2021-06-30 Asher Spector , Lucas Janson

An important problem in machine learning and statistics is to identify features that causally affect the outcome. This is often impossible to do from purely observational data, and a natural relaxation is to identify features that are…

Machine Learning · Statistics 2019-05-30 Jaime Roquero Gimenez , Amirata Ghorbani , James Zou

Variable selection properties of procedures utilizing penalized-likelihood estimates is a central topic in the study of high dimensional linear regression problems. Existing literature emphasizes the quality of ranking of the variables by…

Statistics Theory · Mathematics 2022-04-28 Asaf Weinstein , Weijie J. Su , Małgorzata Bogdan , Rina F. Barber , Emmanuel J. Candès

Identifying which variables do influence a response while controlling false positives pervades statistics and data science. In this paper, we consider a scenario in which we only have access to summary statistics, such as the values of…

Model-X knockoffs is a wrapper that transforms essentially any feature importance measure into a variable selection algorithm, which discovers true effects while rigorously controlling the expected fraction of false positives. A frequently…

Methodology · Statistics 2024-03-12 Stephen Bates , Emmanuel Candès , Lucas Janson , Wenshuo Wang

In this article, we propose a novel strategy for conducting variable selection without prior model topology knowledge using the knockoff method with boosted tree models. Our method is inspired by the original knockoff method, where the…

Methodology · Statistics 2020-02-24 Tao Jiang , Yuanyuan Li , Alison A. Motsinger-Reif

In many scientific fields, researchers are interested in discovering features with substantial effect on the response from a large number of features while controlling the proportion of false discoveries. By incorporating the knockoff…

Methodology · Statistics 2023-02-28 Jiaqi Gu , Guosheng Yin

This paper introduces an innovative method for conducting conditional independence testing in high-dimensional data, facilitating the automated discovery of significant associations within distinct subgroups of a population, all while…

Methodology · Statistics 2023-09-19 Matteo Sesia , Tianshu Sun

While data-driven confounder selection requires careful consideration, it is frequently employed in observational studies. Widely recognized criteria for confounder selection include the minimal-set approach, which involves selecting…

Methodology · Statistics 2025-08-21 Kazuharu Harada , Masataka Taguri

Feature selection prepares the AI-readiness of data by eliminating redundant features. Prior research falls into two primary categories: i) Supervised Feature Selection, which identifies the optimal feature subset based on their relevance…

Machine Learning · Computer Science 2024-03-08 Xinyuan Wang , Dongjie Wang , Wangyang Ying , Rui Xie , Haifeng Chen , Yanjie Fu

This paper introduces a machine for sampling approximate model-X knockoffs for arbitrary and unspecified data distributions using deep generative models. The main idea is to iteratively refine a knockoff sampling mechanism until a criterion…

Methodology · Statistics 2020-03-03 Yaniv Romano , Matteo Sesia , Emmanuel J. Candès