中文
相关论文

相关论文: One-at-a-time knockoffs: controlled false discover…

200 篇论文

We apply the knockoff procedure to factor selection in finance. By building fake but realistic factors, this procedure makes it possible to control the fraction of false discovery in a given set of factors. To show its versatility, we apply…

统计金融 · 定量金融 2021-07-07 Damien Challet , Christian Bongiorno , Guillaume Pelletier

Although sparse autoencoders (SAEs) are crucial for identifying interpretable features in neural networks, it is still challenging to distinguish between real computational patterns and erroneous correlations. We introduce Model-X knockoffs…

机器学习 · 计算机科学 2025-11-18 Tsogt-Ochir Enkhbayar

Variable selection plays a crucial role in enhancing modeling effectiveness across diverse fields, addressing the challenges posed by high-dimensional datasets of correlated variables. This work introduces a novel approach namely Knockoff…

机器学习 · 统计学 2025-01-31 Xiaochen Zhang , Yunfeng Cai , Haoyi Xiong

We propose the use of a new false discovery rate (FDR) controlling procedure as a model selection penalized method, and compare its performance to that of other penalized methods over a wide range of realistic settings: nonorthogonal design…

应用统计 · 统计学 2009-05-19 Yoav Benjamini , Yulia Gavrilov

With the rapid growth of crowdsourcing platforms it has become easy and relatively inexpensive to collect a dataset labeled by multiple annotators in a short time. However due to the lack of control over the quality of the annotators, some…

机器学习 · 统计学 2016-06-17 Qianqian Xu , Jiechao Xiong , Xiaochun Cao , Yuan Yao

This paper is concerned with false discovery rate (FDR) control in large-scale multiple testing problems. We first propose a new data-driven testing procedure for controlling the FDR in large-scale t-tests for one-sample mean problem. The…

统计理论 · 数学 2020-03-02 Changliang Zou , Haojie Ren , Xu Guo , Runze Li

This paper proposes a model-free and data-adaptive feature screening method for ultra-high dimensional datasets. The proposed method is based on the projection correlation which measures the dependence between two random vectors. This…

统计方法学 · 统计学 2021-02-16 Wanjun Liu , Yuan Ke , Jingyuan Liu , Runze Li

Conditional testing via the knockoff framework allows one to identify -- among large number of possible explanatory variables -- those that carry unique information about an outcome of interest, and also provides a false discovery rate…

统计方法学 · 统计学 2024-03-05 Benjamin B Chu , Jiaqi Gu , Zhaomeng Chen , Tim Morrison , Emmanuel Candes , Zihuai He , Chiara Sabatti

Recently, a growing body of research has focused on either optimizing CTR model architectures to better model feature interactions or refining training objectives to aid parameter learning, thereby achieving better predictive performance.…

机器学习 · 计算机科学 2026-05-27 Moyu Zhang , Yun Chen , Yujun Jin , Jinxin Hu , Yu Zhang , Xiaoyi Zeng

Genomic data are subject to various sources of confounding, such as demographic variables, biological heterogeneity, and batch effects. To identify genomic features associated with a variable of interest in the presence of confounders, the…

统计方法学 · 统计学 2025-12-08 Asmita Roy , Jun Chen , Xianyang Zhang

We consider the variable selection problem, which seeks to identify important variables influencing a response $Y$ out of many candidate features $X_1, \ldots, X_p$. We wish to do so while offering finite-sample guarantees about the…

统计方法学 · 统计学 2019-02-12 Rina Foygel Barber , Emmanuel J. Candès , Richard J. Samworth

Controlling the false discovery rate (FDR) is a powerful approach to multiple testing. In many applications, the tested hypotheses have an inherent hierarchical structure. In this paper, we focus on the fixed sequence structure where the…

统计方法学 · 统计学 2016-11-11 Gavin Lynch , Wenge Guo , Sanat K. Sarkar , Helmut Finner

We consider the problem of variable selection in regression models. In particular, we are interested in selecting explanatory covariates linked with the response variable and we want to determine which covariates are relevant, that is which…

统计方法学 · 统计学 2019-07-09 Anne Gégout-Petit , Aurélie Gueudin-Muller , Clémence Karmann

We develop a new class of distribution--free multiple testing rules for false discovery rate (FDR) control under general dependence. A key element in our proposal is a symmetrized data aggregation (SDA) approach to incorporating the…

统计方法学 · 统计学 2021-05-27 Lilun Du , Xu Guo , Wenguang Sun , Changliang Zou

Large-scale hypothesis testing is central to modern science, where controlling the False Discovery Rate (FDR) has become the standard approach to managing false positives across many simultaneous tests. Hypotheses rarely exist in isolation;…

统计方法学 · 统计学 2026-05-19 Binyamin Perets , Shie Mannor

False discovery rate (FDR) control is a popular approach for maintaining the integrity of statistical analyses, especially in high-dimensional data settings, where multiple comparisons increase the risk of false positives. FDR control has…

信号处理 · 电气工程与系统科学 2026-03-03 Fabian Scheidt , Jasin Machkour , Michael Muma

Feature selection is central to contemporary high-dimensional data analysis. Grouping structure among features arises naturally in various scientific problems. Many methods have been proposed to incorporate the grouping structure…

机器学习 · 计算机科学 2019-05-28 Guangyu Zhu , Tingting Zhao

Stability and reproducibility are essential considerations in various applications of statistical methods. False Discovery Rate (FDR) control methods are able to control false signals in scientific discoveries. However, many FDR control…

统计方法学 · 统计学 2025-12-22 Jiajun Sun , Zhanrui Cai , Wei Zhong

This paper introduces a novel Knockoff-guided compressive sensing framework, referred to as \TheName{}, which enhances signal recovery by leveraging precise false discovery rate (FDR) control during the support identification phase. Unlike…

机器学习 · 统计学 2025-06-02 Xiaochen Zhang , Haoyi Xiong

Identifying which variables do influence a response while controlling false positives pervades statistics and data science. In this paper, we consider a scenario in which we only have access to summary statistics, such as the values of…

统计方法学 · 统计学 2024-02-21 Zhaomeng Chen , Zihuai He , Benjamin B. Chu , Jiaqi Gu , Tim Morrison , Chiara Sabatti , Emmanuel Candès