中文
相关论文

相关论文: Differentially Private Variable Selection via the …

200 篇论文

The Model-X knockoffs is a practical methodology for variable selection, which stands out from other selection strategies since it allows for the control of the false discovery rate (FDR), relying on finite-sample guarantees. In this…

Many contemporary large-scale applications involve building interpretable models linking a large set of potential covariates to a response in a nonlinear fashion, such as when the response is binary. Although this modeling problem has been…

统计方法学 · 统计学 2017-12-13 Emmanuel Candes , Yingying Fan , Lucas Janson , Jinchi Lv

Identifying variants that carry substantial information on the trait of interest remains a core topic in genetic studies. In analyzing the EADB-UKBB dataset to identify genetic variants associated with Alzheimer's disease (AD), however, we…

应用统计 · 统计学 2025-12-04 Jiaqi Gu , Zhaomeng Chen , Zihuai He

The problem of selecting a handful of truly relevant variables in supervised machine learning algorithms is a challenging problem in terms of untestable assumptions that must hold and unavailability of theoretical assurances that selection…

统计方法学 · 统计学 2023-11-10 Mehdi Rostami , Olli Saarela

Sorted L-One Penalized Estimation (SLOPE) has shown the nice theoretical property as well as empirical behavior recently on the false discovery rate (FDR) control of high-dimensional feature selection by adaptively imposing the…

统计理论 · 数学 2023-02-22 Jingxuan Liang , Hong Chen , Xuelin Zhang , Weifu Li , Xin Tang

Power and reproducibility are key to enabling refined scientific discoveries in contemporary big data applications with general high-dimensional nonlinear models. In this paper, we provide theoretical foundations on the power and robustness…

统计理论 · 数学 2017-09-04 Yingying Fan , Emre Demirkaya , Gaorong Li , Jinchi Lv

This paper studies the distributed conditional feature screening for massive data with ultrahigh-dimensional features. Specifically, three distributed partial correlation feature screening methods (SAPS, ACPS and JDPS methods) are firstly…

统计方法学 · 统计学 2024-03-12 Naiwen Pang , Xiaochao Xia

Feature selection is central to contemporary high-dimensional data analysis. Grouping structure among features arises naturally in various scientific problems. Many methods have been proposed to incorporate the grouping structure…

机器学习 · 计算机科学 2019-05-28 Guangyu Zhu , Tingting Zhao

Stability and reproducibility are essential considerations in various applications of statistical methods. False Discovery Rate (FDR) control methods are able to control false signals in scientific discoveries. However, many FDR control…

统计方法学 · 统计学 2025-12-22 Jiajun Sun , Zhanrui Cai , Wei Zhong

Testing multiple hypotheses of conditional independence with provable error rate control is a fundamental problem with various applications. To infer conditional independence with family-wise error rate (FWER) control when only summary…

统计方法学 · 统计学 2023-10-17 Catherine Xinrui Yu , Jiaqi Gu , Zhaomeng Chen , Zihuai He

Knockoff variable selection is a powerful framework that creates synthetic knockoff variables to mirror the correlation structure of the observed features, enabling principled control of the false discovery rate in variable selection.…

统计方法学 · 统计学 2025-08-21 Evan Mason , Zhe Fei

We consider the problem of variable selection in regression models. In particular, we are interested in selecting explanatory covariates linked with the response variable and we want to determine which covariates are relevant, that is which…

统计方法学 · 统计学 2019-07-09 Anne Gégout-Petit , Aurélie Gueudin-Muller , Clémence Karmann

False discovery rate (FDR) control is a popular approach for maintaining the integrity of statistical analyses, especially in high-dimensional data settings, where multiple comparisons increase the risk of false positives. FDR control has…

信号处理 · 电气工程与系统科学 2026-03-03 Fabian Scheidt , Jasin Machkour , Michael Muma

We propose a ranking and selection procedure to prioritize relevant predictors and control false discovery proportion (FDP) of variable selection. Our procedure utilizes a new ranking method built upon the de-sparsified Lasso estimator. We…

统计方法学 · 统计学 2018-12-12 X. Jessie Jeng , Xiongzhi Chen

The concept of $k$-FWER has received much attention lately as an appropriate error rate for multiple testing when one seeks to control at least $k$ false rejections, for some fixed $k\ge 1$. A less conservative notion, the $k$-FDR, has been…

统计理论 · 数学 2009-06-18 Sanat K. Sarkar , Wenge Guo

Although sparse autoencoders (SAEs) are crucial for identifying interpretable features in neural networks, it is still challenging to distinguish between real computational patterns and erroneous correlations. We introduce Model-X knockoffs…

机器学习 · 计算机科学 2025-11-18 Tsogt-Ochir Enkhbayar

Controlling the false discovery rate (FDR) in high-dimensional variable selection requires balancing rigorous error control with statistical power. Existing methods with provable guarantees are often overly conservative, creating a…

统计方法学 · 统计学 2026-02-06 Arnau Vilella , Jasin Machkour , Michael Muma , Daniel P. Palomar

Considering the knockoff-based multiple testing framework of Barber and Cand\`es [2015], we revisit the method of Sarkar and Tang [2022] and identify it as a specific case of an un-normalized e-value weighted Benjamini-Hochberg procedure.…

统计方法学 · 统计学 2026-02-13 Aniket Biswas , Aaditya Ramdas

Feature selection prepares the AI-readiness of data by eliminating redundant features. Prior research falls into two primary categories: i) Supervised Feature Selection, which identifies the optimal feature subset based on their relevance…

机器学习 · 计算机科学 2024-03-08 Xinyuan Wang , Dongjie Wang , Wangyang Ying , Rui Xie , Haifeng Chen , Yanjie Fu

Simultaneously finding multiple influential variables and controlling the false discovery rate (FDR) for linear regression models is a fundamental problem. We here propose the Gaussian Mirror (GM) method, which creates for each predictor…

统计方法学 · 统计学 2021-03-22 Xin Xing , Zhigen Zhao , Jun S. Liu