中文
相关论文

相关论文: Knockoff Boosted Tree for Model-Free Variable Sele…

200 篇论文

We present a new procedure for enhanced variable selection for component-wise gradient boosting. Statistical boosting is a computational approach that emerged from machine learning, which allows to fit regression models in the presence of…

统计方法学 · 统计学 2022-02-04 Annika Strömer , Christian Staerk , Nadja Klein , Leonie Weinhold , Stephanie Titze , Andreas Mayr

In modern scientific research, the objective is often to identify which variables are associated with an outcome among a large class of potential predictors. This goal can be achieved by selecting variables in a manner that controls the the…

统计方法学 · 统计学 2023-10-10 Yushu Shi , Michael Martens

We propose a computationally intensive method, the random lasso method, for variable selection in linear models. The method consists of two major steps. In step 1, the lasso method is applied to many bootstrap samples, each using a set of…

应用统计 · 统计学 2011-04-19 Sijian Wang , Bin Nan , Saharon Rosset , Ji Zhu

The rapid generation of complex, highly skewed, and zero-inflated multi-source count data poses significant challenges for variable selection, particularly in biomedical domains like tumor development and metabolic dysregulation. To address…

应用统计 · 统计学 2025-11-11 Shan Tang , Shanjun Mao , Shourong Ma , Falong Tan

Consider a case-control study in which we have a random sample, constructed in such a way that the proportion of cases in our sample is different from that in the general population---for instance, the sample is constructed to achieve a…

统计方法学 · 统计学 2019-01-01 Rina Foygel Barber , Emmanuel Candes

The Model-X knockoffs is a practical methodology for variable selection, which stands out from other selection strategies since it allows for the control of the false discovery rate (FDR), relying on finite-sample guarantees. In this…

Structured additive distributional copula regression allows to model the joint distribution of multivariate outcomes by relating all distribution parameters to covariates. Estimation via statistical boosting enables accounting for…

Power and reproducibility are key to enabling refined scientific discoveries in contemporary big data applications with general high-dimensional nonlinear models. In this paper, we provide theoretical foundations on the power and robustness…

统计理论 · 数学 2017-09-04 Yingying Fan , Emre Demirkaya , Gaorong Li , Jinchi Lv

The Model-X knockoff procedure has recently emerged as a powerful approach for feature selection with statistical guarantees. The advantage of knockoff is that if we have a good model of the features X, then we can identify salient features…

机器学习 · 统计学 2019-05-30 Jaime Roquero Gimenez , James Zou

Genome-wide association studies (GWAS) often find association signals between many genetic variants and traits of interest in a genomic region. Functional annotations of these variants provide valuable prior information that helps…

统计方法学 · 统计学 2026-01-07 Xiangyu Zhang , Lijun Wang , Changjun Li , Chen Lin , Hongyu Zhao

The fixed-X knockoff filter is a flexible framework for variable selection with false discovery rate (FDR) control in linear models with arbitrary design matrices (of full column rank) and it allows for finite-sample selective inference via…

统计理论 · 数学 2023-11-28 Mehrdad Pournaderi , Yu Xiang

Few problems in statistics are as perplexing as variable selection in the presence of very many redundant covariates. The variable selection problem is most familiar in parametric environments such as the linear model or additive variants…

统计方法学 · 统计学 2021-02-25 Yi Liu , Veronika Ročková , Yuexi Wang

Model-X knockoff has garnered significant attention among various feature selection methods due to its guarantees for controlling the false discovery rate (FDR). Since its introduction in parametric design, knockoff techniques have evolved…

机器学习 · 计算机科学 2024-11-11 Hongyu Shen , Yici Yan , Zhizhen Zhao

High-dimensional variable selection has emerged as one of the prevailing statistical challenges in the big data revolution. Many variable selection methods have been adapted for identifying single nucleotide polymorphisms (SNPs) linked to…

统计方法学 · 统计学 2024-08-21 Justin J. Van Ee , Diana Gamba , Jesse R. Lasky , Megan L. Vahsen , Mevin B. Hooten

We propose a new method to learn the structure of a Gaussian graphical model with finite sample false discovery rate control. Our method builds on the knockoff framework of Barber and Cand\`{e}s for linear models. We extend their approach…

统计方法学 · 统计学 2021-04-20 Jinzhou Li , Marloes H. Maathuis

We consider the problem of assessing the importance of multiple variables or factors from a dataset when side information is available. In principle, using side information can allow the statistician to pay attention to variables with a…

统计方法学 · 统计学 2020-01-23 Zhimei Ren , Emmanuel Candès

Gradient boosting algorithms construct a regression predictor using a linear combination of ``base learners''. Boosting also offers an approach to obtaining robust non-parametric regression estimators that are scalable to applications with…

统计方法学 · 统计学 2020-08-11 Xiaomeng Ju , Matías Salibián-Barrera

Model-X knockoffs allows analysts to perform feature selection using almost any machine learning algorithm while still provably controlling the expected proportion of false discoveries. To apply model-X knockoffs, one must construct…

统计方法学 · 统计学 2021-06-30 Asher Spector , Lucas Janson

The idea of "stratified medicine" is an important driver of methodological research on the identification of predictive biomarkers. Most methods proposed so far for this purpose have been developed for the use on randomized data only.…

统计方法学 · 统计学 2022-12-19 Julia Krzykalla , Axel Benner , Annette Kopp-Schneider

The instability in the selection of models is a major concern with data sets containing a large number of covariates. This paper deals with variable selection methodology in the case of high-dimensional problems where the response variable…

应用统计 · 统计学 2012-03-23 Marie Walschaerts , Eve Leconte , Philippe Besse