中文
相关论文

相关论文: Knockoffs for the mass: new feature importance sta…

200 篇论文

Modern scientific studies often require the identification of a subset of relevant explanatory variables, in the attempt to understand an interesting phenomenon. Several statistical methods have been developed to automate this task, but…

统计方法学 · 统计学 2019-05-14 Matteo Sesia , Chiara Sabatti , Emmanuel J. Candès

This paper develops a framework for testing for associations in a possibly high-dimensional linear model where the number of features/variables may far exceed the number of observational units. In this framework, the observations are split…

统计方法学 · 统计学 2018-05-04 Rina Foygel Barber , Emmanuel J. Candes

The false discovery rate (FDR)---the expected fraction of spurious discoveries among all the discoveries---provides a popular statistical assessment of the reproducibility of scientific studies in various disciplines. In this work, we…

机器学习 · 统计学 2015-11-10 Weijie Su , Junyang Qian , Linxi Liu

This paper introduces a machine for sampling approximate model-X knockoffs for arbitrary and unspecified data distributions using deep generative models. The main idea is to iteratively refine a knockoff sampling mechanism until a criterion…

统计方法学 · 统计学 2020-03-03 Yaniv Romano , Matteo Sesia , Emmanuel J. Candès

This paper develops a method based on model-X knockoffs to find conditional associations that are consistent across diverse environments, controlling the false discovery rate. The motivation for this problem is that large data sets may…

统计方法学 · 统计学 2021-06-09 Shuangning Li , Matteo Sesia , Yaniv Romano , Emmanuel Candès , Chiara Sabatti

Testing multiple hypotheses of conditional independence with provable error rate control is a fundamental problem with various applications. To infer conditional independence with family-wise error rate (FWER) control when only summary…

统计方法学 · 统计学 2023-10-17 Catherine Xinrui Yu , Jiaqi Gu , Zhaomeng Chen , Zihuai He

Model-X knockoffs is a general procedure that can leverage any feature importance measure to produce a variable selection algorithm, which discovers true effects while rigorously controlling the number or fraction of false positives.…

统计方法学 · 统计学 2020-12-07 Zhimei Ren , Yuting Wei , Emmanuel Candès

Thanks to its fine balance between model flexibility and interpretability, the nonparametric additive model has been widely used, and variable selection for this type of model has been frequently studied. However, none of the existing…

统计方法学 · 统计学 2022-01-10 Xiaowu Dai , Xiang Lyu , Lexin Li

This paper introduces an innovative method for conducting conditional independence testing in high-dimensional data, facilitating the automated discovery of significant associations within distinct subgroups of a population, all while…

统计方法学 · 统计学 2023-09-19 Matteo Sesia , Tianshu Sun

In 2015, Barber and Candes introduced a new variable selection procedure called the knockoff filter to control the false discovery rate (FDR) and prove that this method achieves exact FDR control. Inspired by the work of Barber and Candes…

统计方法学 · 统计学 2019-07-23 Jiajie Chen , Anthony Hou , Thomas Y. Hou

The knockoff filter is a recent false discovery rate (FDR) control method for high-dimensional linear models. We point out that knockoff has three key components: ranking algorithm, augmented design, and symmetric statistic, and each…

统计理论 · 数学 2024-02-14 Zheng Tracy Ke , Jun S. Liu , Yucong Ma

In this article, we propose a novel strategy for conducting variable selection without prior model topology knowledge using the knockoff method with boosted tree models. Our method is inspired by the original knockoff method, where the…

统计方法学 · 统计学 2020-02-24 Tao Jiang , Yuanyuan Li , Alison A. Motsinger-Reif

Knockoffs provide a general framework for controlling the false discovery rate when performing variable selection. Much of the Knockoffs literature focuses on theoretical challenges and we recognize a need for bringing some of the current…

统计方法学 · 统计学 2020-10-28 Matthias Kormaksson , Luke J. Kelly , Xuan Zhu , Sibylle Haemmerle , Luminita Pricop , David Ohlssen

We describe a series of algorithms that efficiently implement Gaussian model-X knockoffs to control the false discovery rate on large scale feature selection problems. Identifying the knockoff distribution requires solving a large scale…

机器学习 · 计算机科学 2020-06-17 Armin Askari , Quentin Rebjock , Alexandre d'Aspremont , Laurent El Ghaoui

Continuous improvement in medical imaging techniques allows the acquisition of higher-resolution images. When these are used in a predictive setting, a greater number of explanatory variables are potentially related to the dependent…

统计理论 · 数学 2019-03-13 Tuan-Binh Nguyen , Jérôme-Alexis Chevalier , Bertrand Thirion

Feature selection has remained a daunting challenge in machine learning and artificial intelligence, where increasingly complex, high-dimensional datasets demand principled strategies for isolating the most informative predictors. Despite…

机器学习 · 统计学 2025-12-02 Mousam Sinha , Tirtha Sarathi Ghosh , Ridam Pal

Variable selection plays a crucial role in enhancing modeling effectiveness across diverse fields, addressing the challenges posed by high-dimensional datasets of correlated variables. This work introduces a novel approach namely Knockoff…

机器学习 · 统计学 2025-01-31 Xiaochen Zhang , Yunfeng Cai , Haoyi Xiong

Many contemporary large-scale applications involve building interpretable models linking a large set of potential covariates to a response in a nonlinear fashion, such as when the response is binary. Although this modeling problem has been…

统计方法学 · 统计学 2017-12-13 Emmanuel Candes , Yingying Fan , Lucas Janson , Jinchi Lv

We introduce local conditional hypotheses that express how the relation between explanatory variables and outcomes changes across different contexts, described by covariates. By expanding upon the model-X knockoff filter, we show how to…

统计方法学 · 统计学 2026-01-12 Paula Gablenz , Matteo Sesia , Tianshu Sun , Chiara Sabatti

Although there is a huge literature on feature selection for the Cox model, none of the existing approaches can control the false discovery rate (FDR) unless the sample size tends to infinity. In addition, there is no formal power analysis…

统计方法学 · 统计学 2023-08-02 Daoji Li , Jinzhao Yu , Hui Zhao