中文
相关论文

相关论文: Panning for Gold: Model-X Knockoffs for High-dimen…

200 篇论文

The recent proliferation of high-dimensional data, such as electronic health records and genetics data, offers new opportunities to find novel predictors of outcomes. Presented with a large set of candidate features, interest often lies in…

统计方法学 · 统计学 2024-09-24 Michael J. Martens , Anjishnu Banerjee , Xinran Qi , Yushu Shi

We tackle the problem of selecting from among a large number of variables those that are 'important' for an outcome. We consider situations where groups of variables are also of interest in their own right. For example, each variable might…

统计方法学 · 统计学 2018-08-13 Eugene Katsevich , Chiara Sabatti

In modern scientific research, the objective is often to identify which variables are associated with an outcome among a large class of potential predictors. This goal can be achieved by selecting variables in a manner that controls the the…

统计方法学 · 统计学 2023-10-10 Yushu Shi , Michael Martens

Selecting important features in high-dimensional survival analysis is critical for identifying confirmatory biomarkers while maintaining rigorous error control. In this paper, we propose a derandomized knockoffs procedure for Cox regression…

统计方法学 · 统计学 2025-12-15 Rui Liu , Nan Sun

The Model-X knockoffs is a practical methodology for variable selection, which stands out from other selection strategies since it allows for the control of the false discovery rate (FDR), relying on finite-sample guarantees. In this…

The knockoff filter of Barber and Candes (arXiv:1404.5609) is a flexible framework for multiple testing in supervised learning models, based on introducing synthetic predictor variables to control the false discovery rate (FDR). Using the…

统计方法学 · 统计学 2024-11-26 Yixiang Luo , William Fithian , Lihua Lei

Variable selection has been widely used in data analysis for the past decades, and it becomes increasingly important in the Big Data era as there are usually hundreds of variables available in a dataset. To enhance interpretability of a…

统计方法学 · 统计学 2020-08-17 Yuxiang Xie , Kwun Chuen Gary Chan

Predictive modeling often uses black box machine learning methods, such as deep neural networks, to achieve state-of-the-art performance. In scientific domains, the scientist often wishes to discover which features are actually important…

机器学习 · 统计学 2020-08-03 Mukund Sudarshan , Wesley Tansey , Rajesh Ranganath

We consider the problem of assessing the importance of multiple variables or factors from a dataset when side information is available. In principle, using side information can allow the statistician to pay attention to variables with a…

统计方法学 · 统计学 2020-01-23 Zhimei Ren , Emmanuel Candès

An important problem in machine learning and statistics is to identify features that causally affect the outcome. This is often impossible to do from purely observational data, and a natural relaxation is to identify features that are…

机器学习 · 统计学 2019-05-30 Jaime Roquero Gimenez , Amirata Ghorbani , James Zou

Let $X=(X_1,\ldots,X_p)$ be a $p$-variate random vector and $F$ a fixed finite set. In a number of applications, mainly in genetics, it turns out that $X_i\in F$ for each $i=1,\ldots,p$. Despite the latter fact, to obtain a knockoff…

统计理论 · 数学 2024-10-15 Emanuela Dreassi , Luca Pratelli , Pietro Rigo

Interpretability and stability are two important features that are desired in many contemporary big data applications arising in economics and finance. While the former is enjoyed to some extent by many existing forecasting approaches, the…

统计理论 · 数学 2018-09-14 Yingying Fan , Jinchi Lv , Mahrad Sharifvaghefi , Yoshimasa Uematsu

In many scientific problems, researchers try to relate a response variable $Y$ to a set of potential explanatory variables $X = (X_1,\dots,X_p)$, and start by trying to identify variables that contribute to this relationship. In statistical…

统计理论 · 数学 2020-10-07 Wenshuo Wang , Lucas Janson

As new Model-X knockoff construction techniques are developed, primarily concerned with determining the correct conditional distribution from which to sample, we focus less on deriving the correct multivariate distribution and instead ask…

统计方法学 · 统计学 2025-02-05 Christopher Hemmens , Stephan Robert-Nicoud

Variable selection plays a crucial role in enhancing modeling effectiveness across diverse fields, addressing the challenges posed by high-dimensional datasets of correlated variables. This work introduces a novel approach namely Knockoff…

机器学习 · 统计学 2025-01-31 Xiaochen Zhang , Yunfeng Cai , Haoyi Xiong

We present a novel method for controlling the $k$-familywise error rate ($k$-FWER) in the linear regression setting using the knockoffs framework first introduced by Barber and Cand\`es. Our procedure, which we also refer to as knockoffs,…

统计方法学 · 统计学 2015-11-10 Lucas Janson , Weijie Su

Controlled feature selection aims to discover the features a response depends on while limiting the false discovery rate (FDR) to a predefined level. Recently, multiple deep-learning-based methods have been proposed to perform controlled…

机器学习 · 统计学 2022-10-24 Derek Hansen , Brian Manzo , Jeffrey Regier

Controlling false discovery rate (FDR) is crucial for variable selection, multiple testing, among other signal detection problems. In literature, there is certainly no shortage of FDR control strategies when selecting individual features,…

统计方法学 · 统计学 2022-04-11 Jingyuan Liu , Ao Sun , Yuan Ke

We make some initial attempt to establish the theoretical and methodological foundation for the model-X knockoffs inference for time series data. We suggest the method of time series knockoffs inference (TSKI) by exploiting the ideas of…

统计方法学 · 统计学 2025-03-03 Chien-Ming Chi , Yingying Fan , Ching-Kang Ing , Jinchi Lv

Continuous improvement in medical imaging techniques allows the acquisition of higher-resolution images. When these are used in a predictive setting, a greater number of explanatory variables are potentially related to the dependent…

统计理论 · 数学 2019-03-13 Tuan-Binh Nguyen , Jérôme-Alexis Chevalier , Bertrand Thirion