中文
相关论文

相关论文: Knockoff Boosted Tree for Model-Free Variable Sele…

200 篇论文

Controlling false discovery rate (FDR) is crucial for variable selection, multiple testing, among other signal detection problems. In literature, there is certainly no shortage of FDR control strategies when selecting individual features,…

统计方法学 · 统计学 2022-04-11 Jingyuan Liu , Ao Sun , Yuan Ke

Among the most popular variable selection procedures in high-dimensional regression, Lasso provides a solution path to rank the variables and determines a cut-off position on the path to select variables and estimate coefficients. In this…

统计方法学 · 统计学 2018-06-19 X. Jessie Jeng , Huimin Peng , Wenbin Lu

Decision trees and their ensembles are endowed with a rich set of diagnostic tools for ranking and screening variables in a predictive model. Despite the widespread use of tree based variable importance measures, pinning down their…

机器学习 · 统计学 2020-12-14 Jason M. Klusowski , Peter M. Tian

Controlled variable selection is an important analytical step in various scientific fields, such as brain imaging or genomics. In these high-dimensional data settings, considering too many variables leads to poor models and high costs,…

统计方法学 · 统计学 2023-10-17 Alexandre Blain , Bertrand Thirion , Olivier Grisel , Pierre Neuvial

In this paper we deepen and enlarge the reflection on the possible advantages of a knockoff approach to genome wide association studies (Sesia et al., 2018), starting from the discussions in Bottolo & Richardson (2019); Jewell & Witten…

统计方法学 · 统计学 2019-03-15 Matteo Sesia , Chiara Sabatti , Emmanuel J. Candès

The goal of feature selection is to identify important features that are relevant to explain an outcome variable. Most of the work in this domain has focused on identifying globally relevant features, which are features that are related to…

机器学习 · 统计学 2019-05-30 Jaime Roquero Gimenez , James Zou

Random forests are a statistical learning technique that use bootstrap aggregation to average high-variance and low-bias trees. Improvements to random forests, such as applying Lasso regression to the tree predictions, have been proposed in…

机器学习 · 统计学 2025-11-13 Jing Shang , James Bannon , Benjamin Haibe-Kains , Robert Tibshirani

Tree-boosting is a widely used machine learning technique for tabular data. However, its out-of-sample accuracy is critically dependent on multiple hyperparameters. In this article, we empirically compare several popular methods for…

机器学习 · 计算机科学 2026-05-29 Floris Jan Koster , Fabio Sigrist

We consider the problem of identifying significant predictors in large data bases, where the response variable depends on the linear combination of explanatory variables through an unknown link function, corrupted with the noise from the…

统计方法学 · 统计学 2019-11-19 Wojciech Rejchel , Malgorzata Bogdan

Multiple comparisons in hypothesis testing often encounter structural constraints in various applications. For instance, in structural Magnetic Resonance Imaging for Alzheimer's Disease, the focus extends beyond examining atrophic brain…

统计方法学 · 统计学 2023-11-08 Yang Cao , Xinwei Sun , Yuan Yao

High-dimensional variable selection in the proportional hazards (PH) model has many successful applications in different areas. In practice, data may involve confounding variables that do not satisfy the PH assumption, in which case the…

统计计算 · 统计学 2018-03-22 Emily Morris , Kevin He , Yanming Li , Yi Li , Jian Kang

Molecular data from tumor profiles is high dimensional. Tumor profiles can be characterized by tens of thousands of gene expression features. Due to the size of the gene expression feature set machine learning methods are exposed to noisy…

机器学习 · 计算机科学 2020-07-14 Martin Palazzo , Pierre Beauseroy , Patricio Yankilevich

We propose the group knockoff filter, a method for false discovery rate control in a linear regression setting where the features are grouped, and we would like to select a set of relevant groups which have a nonzero effect on the response.…

统计方法学 · 统计学 2016-02-12 Ran Dai , Rina Foygel Barber

In this article we propose a boosting algorithm for regression with functional explanatory variables and scalar responses. The algorithm uses decision trees constructed with multiple projections as the "base-learners", which we call…

统计方法学 · 统计学 2023-04-07 Xiaomeng Ju , Matías Salibián-Barrera

We propose a Bayesian variable selection method in the framework of modal regression for heavy-tailed responses. An efficient expectation-maximization algorithm is employed to expedite parameter estimation. A test statistic is constructed…

统计方法学 · 统计学 2025-10-29 Jiasong Duan , Hongmei Zhang , Xianzheng Huang

Conditional testing via the knockoff framework allows one to identify -- among large number of possible explanatory variables -- those that carry unique information about an outcome of interest, and also provides a false discovery rate…

统计方法学 · 统计学 2024-03-05 Benjamin B Chu , Jiaqi Gu , Zhaomeng Chen , Tim Morrison , Emmanuel Candes , Zihuai He , Chiara Sabatti

Modern statistical analyses often involve testing large numbers of hypotheses. In many situations, these hypotheses may have an underlying tree structure that not only helps determine the order that tests should be conducted but also…

统计方法学 · 统计学 2019-03-19 Yunxiao Li , Yi-Juan Hu , Glen A. Satten

This paper proposes a reliable neural network pruning algorithm by setting up a scientific control. Existing pruning methods have developed various hypotheses to approximate the importance of filters to the network and then execute filter…

计算机视觉与模式识别 · 计算机科学 2021-01-12 Yehui Tang , Yunhe Wang , Yixing Xu , Dacheng Tao , Chunjing Xu , Chao Xu , Chang Xu

In the fight against hard-to-treat diseases such as cancer, it is often difficult to discover new treatments that benefit all subjects. For regulatory agency approval, it is more practical to identify subgroups of subjects for whom the…

统计方法学 · 统计学 2014-10-09 Wei-Yin Loh , Xu He , Michael Man

It is increasingly common clinically for cancer specimens to be examined using techniques that identify somatic mutations. In principle these mutational profiles can be used to diagnose the tissue of origin, a critical task for the 3-5% of…

统计方法学 · 统计学 2020-07-14 Saptarshi Chakraborty , Colin B. Begg , Ronglai Shen