English
Related papers

Related papers: Gold after Randomized Sand: Model-X Split Knockoff…

200 papers

The knockoff filter is a recent false discovery rate (FDR) control method for high-dimensional linear models. We point out that knockoff has three key components: ranking algorithm, augmented design, and symmetric statistic, and each…

Statistics Theory · Mathematics 2024-02-14 Zheng Tracy Ke , Jun S. Liu , Yucong Ma

Model-X knockoffs is a general procedure that can leverage any feature importance measure to produce a variable selection algorithm, which discovers true effects while rigorously controlling the number or fraction of false positives.…

Methodology · Statistics 2020-12-07 Zhimei Ren , Yuting Wei , Emmanuel Candès

The fixed-X knockoff filter is a flexible framework for variable selection with false discovery rate (FDR) control in linear models with arbitrary design matrices (of full column rank) and it allows for finite-sample selective inference via…

Statistics Theory · Mathematics 2023-11-28 Mehrdad Pournaderi , Yu Xiang

Multiple comparisons in hypothesis testing often encounter structural constraints in various applications. For instance, in structural Magnetic Resonance Imaging for Alzheimer's Disease, the focus extends beyond examining atrophic brain…

Methodology · Statistics 2023-11-08 Yang Cao , Xinwei Sun , Yuan Yao

We propose a unified theoretical framework for studying the robustness of the model-X knockoffs framework by investigating the asymptotic false discovery rate (FDR) control of the practically implemented approximate knockoffs procedure.…

Machine Learning · Statistics 2025-02-11 Yingying Fan , Lan Gao , Jinchi Lv , Xiaocong Xu

Model-X knockoffs is a wrapper that transforms essentially any feature importance measure into a variable selection algorithm, which discovers true effects while rigorously controlling the expected fraction of false positives. A frequently…

Methodology · Statistics 2024-03-12 Stephen Bates , Emmanuel Candès , Lucas Janson , Wenshuo Wang

We consider the variable selection problem, which seeks to identify important variables influencing a response $Y$ out of many candidate features $X_1, \ldots, X_p$. We wish to do so while offering finite-sample guarantees about the…

Methodology · Statistics 2019-02-12 Rina Foygel Barber , Emmanuel J. Candès , Richard J. Samworth

Selecting relevant features associated with a given response variable is an important issue in many scientific fields. Quantifying quality and uncertainty of a selection result via false discovery rate (FDR) control has been of recent…

Methodology · Statistics 2020-12-17 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

This paper develops a framework for testing for associations in a possibly high-dimensional linear model where the number of features/variables may far exceed the number of observational units. In this framework, the observations are split…

Methodology · Statistics 2018-05-04 Rina Foygel Barber , Emmanuel J. Candes

In modern scientific research, the objective is often to identify which variables are associated with an outcome among a large class of potential predictors. This goal can be achieved by selecting variables in a manner that controls the the…

Methodology · Statistics 2023-10-10 Yushu Shi , Michael Martens

Simultaneously performing variable selection and inference in high-dimensional regression models is an open challenge in statistics and machine learning. The increasing availability of vast amounts of variables requires the adoption of…

Methodology · Statistics 2025-05-08 Marco Molinari , Magne Thoresen

Algorithms that ensure reproducible findings from large-scale, high-dimensional data are pivotal in numerous signal processing applications. In recent years, multivariate false discovery rate (FDR) controlling methods have emerged,…

Methodology · Statistics 2024-01-31 Jasin Machkour , Michael Muma , Daniel P. Palomar

False discovery rate (FDR) controlling procedures provide important statistical guarantees for the replicability in signal identification based on multiple hypotheses testing. In many fields of study, FDR controlling procedures are used in…

Methodology · Statistics 2022-10-04 Ran Dai , Cheng Zheng

We investigate the robustness of the model-X knockoffs framework with respect to the misspecified or estimated feature distribution. We achieve such a goal by theoretically studying the feature selection performance of a practically…

Methodology · Statistics 2024-06-06 Yingying Fan , Lan Gao , Jinchi Lv

The recent proliferation of high-dimensional data, such as electronic health records and genetics data, offers new opportunities to find novel predictors of outcomes. Presented with a large set of candidate features, interest often lies in…

Methodology · Statistics 2024-09-24 Michael J. Martens , Anjishnu Banerjee , Xinran Qi , Yushu Shi

We introduce a novel privatization framework for high-dimensional controlled variable selection. Our framework enables rigorous False Discovery Rate (FDR) control under differential privacy constraints. While the Model-X knockoff procedure…

Machine Learning · Statistics 2025-08-08 Yuxuan Tao , Adel Javanmard

This paper proposes a model-free and data-adaptive feature screening method for ultra-high dimensional datasets. The proposed method is based on the projection correlation which measures the dependence between two random vectors. This…

Methodology · Statistics 2021-02-16 Wanjun Liu , Yuan Ke , Jingyuan Liu , Runze Li

Power and reproducibility are key to enabling refined scientific discoveries in contemporary big data applications with general high-dimensional nonlinear models. In this paper, we provide theoretical foundations on the power and robustness…

Statistics Theory · Mathematics 2017-09-04 Yingying Fan , Emre Demirkaya , Gaorong Li , Jinchi Lv

We develop a new class of distribution--free multiple testing rules for false discovery rate (FDR) control under general dependence. A key element in our proposal is a symmetrized data aggregation (SDA) approach to incorporating the…

Methodology · Statistics 2021-05-27 Lilun Du , Xu Guo , Wenguang Sun , Changliang Zou

This paper introduces a machine for sampling approximate model-X knockoffs for arbitrary and unspecified data distributions using deep generative models. The main idea is to iteratively refine a knockoff sampling mechanism until a criterion…

Methodology · Statistics 2020-03-03 Yaniv Romano , Matteo Sesia , Emmanuel J. Candès