English
Related papers

Related papers: A Pseudo Knockoff Filter for Correlated Features

200 papers

This paper studies the distributed conditional feature screening for massive data with ultrahigh-dimensional features. Specifically, three distributed partial correlation feature screening methods (SAPS, ACPS and JDPS methods) are firstly…

Methodology · Statistics 2024-03-12 Naiwen Pang , Xiaochao Xia

Researchers in biomedical studies often work with samples that are not selected uniformly at random from the population of interest, a major example being a case-control study. While these designs are motivated by specific scientific…

This paper presents a survey on some recent advances for the type I error rate control in multiple testing methodology. We consider the problem of controlling the $k$-family-wise error rate (kFWER, probability to make $k$ false discoveries…

Methodology · Statistics 2011-03-15 Etienne Roquain

We develop an extension of the Knockoff Inference procedure, introduced by Barber and Candes (2015). This new method, called Aggregation of Multiple Knockoffs (AKO), addresses the instability inherent to the random nature of Knockoff-based…

Statistics Theory · Mathematics 2020-06-26 Tuan-Binh Nguyen , Jérôme-Alexis Chevalier , Bertrand Thirion , Sylvain Arlot

In the FDR-controlling literature, mirror statistics offer a flexible alternative to $p$-value based procedures. When prior information is available, however, it is unclear how to incorporate mirror statistics in a principled way, and the…

Methodology · Statistics 2026-04-22 Yuanchuan Guo , Buyu Lin , Jun S. Liu

Stability and reproducibility are essential considerations in various applications of statistical methods. False Discovery Rate (FDR) control methods are able to control false signals in scientific discoveries. However, many FDR control…

Methodology · Statistics 2025-12-22 Jiajun Sun , Zhanrui Cai , Wei Zhong

The goal of feature selection is to identify important features that are relevant to explain an outcome variable. Most of the work in this domain has focused on identifying globally relevant features, which are features that are related to…

Machine Learning · Statistics 2019-05-30 Jaime Roquero Gimenez , James Zou

In feature selection problems, knockoffs are synthetic controls for the original features. Employing knockoffs allows analysts to use nearly any variable importance measure or "feature statistic" to select features while rigorously…

Methodology · Statistics 2024-10-02 Asher Spector , William Fithian

Balancing false discovery rate (FDR) control with high statistical power remains a central challenge in high-dimensional variable selection. While several FDR-controlling methods have been proposed, many degrade the original data -- by…

Methodology · Statistics 2025-07-16 Changhu Wang , Ziheng Zhang , Jingyi Jessica Li

We address challenges in variable selection with highly correlated data that are frequently present in finance, economics, but also in complex natural systems as e.g. weather. We develop a robustified version of the knockoff framework,…

Econometrics · Economics 2022-06-14 Konstantin Görgen , Abdolreza Nazemi , Melanie Schienle

One challenge in exploratory association studies using observational data is that the associations between the predictors and the outcome are potentially weak and rare, and the candidate predictors have complex correlation structures. False…

Methodology · Statistics 2025-01-30 Runqiu Wang , Ran Dai , Hongying Dai , Evan French , Cheng Zheng

Model-X knockoffs is a general procedure that can leverage any feature importance measure to produce a variable selection algorithm, which discovers true effects while rigorously controlling the number or fraction of false positives.…

Methodology · Statistics 2020-12-07 Zhimei Ren , Yuting Wei , Emmanuel Candès

Most link prediction methods return estimates of the connection probability of missing edges in a graph. Such output can be used to rank the missing edges from most to least likely to be a true edge, but does not directly provide a…

Methodology · Statistics 2024-03-26 Ariane Marandon

The generalized linear models (GLM) have been widely used in practice to model non-Gaussian response variables. When the number of explanatory features is relatively large, scientific researchers are of interest to perform controlled…

Methodology · Statistics 2020-07-03 Chenguang Dai , Buyu Lin , Xin Xing , Jun S. Liu

Benjamini and Hochberg (1995) proposed the false discovery rate (FDR) as an alternative to the family-wise error rate in multiple testing problems, and proposed a procedure to control the FDR. For discrete data this procedure may be highly…

Methodology · Statistics 2014-01-28 Ruth Heller , Hadas Gur

This paper introduces an innovative method for conducting conditional independence testing in high-dimensional data, facilitating the automated discovery of significant associations within distinct subgroups of a population, all while…

Methodology · Statistics 2023-09-19 Matteo Sesia , Tianshu Sun

This paper proposes a reliable neural network pruning algorithm by setting up a scientific control. Existing pruning methods have developed various hypotheses to approximate the importance of filters to the network and then execute filter…

Computer Vision and Pattern Recognition · Computer Science 2021-01-12 Yehui Tang , Yunhe Wang , Yixing Xu , Dacheng Tao , Chunjing Xu , Chao Xu , Chang Xu

Multiple hypothesis testing is a fundamental problem in high dimensional inference, with wide applications in many scientific fields. In genome-wide association studies, tens of thousands of tests are performed simultaneously to find if any…

Methodology · Statistics 2011-11-16 Jianqing Fan , Xu Han , Weijie Gu

Power and reproducibility are key to enabling refined scientific discoveries in contemporary big data applications with general high-dimensional nonlinear models. In this paper, we provide theoretical foundations on the power and robustness…

Statistics Theory · Mathematics 2017-09-04 Yingying Fan , Emre Demirkaya , Gaorong Li , Jinchi Lv

This paper revisits the following open question in simultaneous testing of multivariate normal means against two-sided alternatives: Can the method of Benjamini and Hochberg (BH, 1995) control the false discovery rate (FDR) without imposing…

Statistics Theory · Mathematics 2023-04-12 Sanat K. Sarkar
‹ Prev 1 3 4 5 6 7 10 Next ›