English
Related papers

Related papers: Setwise Hierarchical Variable Selection and the Ge…

200 papers

There is a challenge in selecting high-dimensional mediators when the mediators have complex correlation structures and interactions. In this work, we frame the high-dimensional mediator selection problem into a series of hypothesis tests…

Methodology · Statistics 2025-09-16 Runqiu Wang , Ran Dai , Jieqiong Wang , Kah Meng Soh , Ziyang Xu , Mohamed Azzam , Hongying Dai , Cheng Zheng

A new statistical procedure (Model-X \cite{candes2018}) has provided a way to identify important factors using any supervised learning method controlling for FDR. This line of research has shown great potential to expand the horizon of…

Methodology · Statistics 2018-10-01 Ying Liu , Cheng Zheng

Biological research often involves testing a growing number of null hypotheses as new data is accumulated over time. We study the problem of online control of the familywise error rate (FWER), that is testing an apriori unbounded sequence…

Methodology · Statistics 2020-03-10 Jinjin Tian , Aaditya Ramdas

Variable selection on the large-scale networks has been extensively studied in the literature. While most of the existing methods are limited to the local functionals especially the graph edges, this paper focuses on selecting the discrete…

Methodology · Statistics 2023-09-18 Lu Zhang , Junwei Lu

We propose the group knockoff filter, a method for false discovery rate control in a linear regression setting where the features are grouped, and we would like to select a set of relevant groups which have a nonzero effect on the response.…

Methodology · Statistics 2016-02-12 Ran Dai , Rina Foygel Barber

Variable selection plays a fundamental role in high-dimensional data analysis. Various methods have been developed for variable selection in recent years. Well-known examples are forward stepwise regression (FSR) and least angle regression…

Methodology · Statistics 2018-02-01 Siliang Gong , Kai Zhang , Yufeng Liu

In this paper, a noisy version of the stochastic block model (NSBM) is introduced and we investigate the three following statistical inferences in this model: estimation of the model parameters, clustering of the nodes and identification of…

Statistics Theory · Mathematics 2019-07-25 Tabea Rebafka , Etienne Roquain , Fanny Villers

With the development of data collection techniques, analysis with a survival response and high-dimensional covariates has become routine. Here we consider an interaction model, which includes a set of low-dimensional covariates, a set of…

Methodology · Statistics 2023-11-27 Weijuan Liang , Qingzhao Zhang , Shuangge Ma

Controlled variable selection is an important analytical step in various scientific fields, such as brain imaging or genomics. In these high-dimensional data settings, considering too many variables leads to poor models and high costs,…

Methodology · Statistics 2023-10-17 Alexandre Blain , Bertrand Thirion , Olivier Grisel , Pierre Neuvial

In modern scientific experiments, we frequently encounter data that have large dimensions, and in some experiments, such high dimensional data arrive sequentially rather than full data being available all at a time. We develop multiple…

Methodology · Statistics 2023-06-09 Rahul Roy , Shyamal K. De , Subir Kumar Bhandari

Recently, Barber and Cand\`es laid the theoretical foundation for a general framework for false discovery rate (FDR) control based on the notion of "knockoffs." A closely related FDR control methodology has long been employed in the…

Methodology · Statistics 2022-03-15 Dong Luo , Arya Ebadi , Yilun He , Kristen Emery , William Stafford Noble , Uri Keich

Obtaining high-quality labels for large datasets is expensive, requiring massive annotations from human experts. While AI models offer a cost-effective alternative by predicting labels, their label quality is compromised by the unavoidable…

Machine Learning · Computer Science 2026-02-17 Huipeng Huang , Wenbo Liao , Huajun Xi , Hao Zeng , Mengchen Zhao , Hongxin Wei

In this paper we introduce and investigate a new rejection curve for asymptotic control of the false discovery rate (FDR) in multiple hypotheses testing problems. We first give a heuristic motivation for this new curve and propose some…

Statistics Theory · Mathematics 2009-03-31 Helmut Finner , Thorsten Dickhaus , Markus Roters

Recently, the scheme of model-X knockoffs was proposed as a promising solution to address controlled feature selection under high-dimensional finite-sample settings. However, the procedure of model-X knockoffs depends heavily on the…

Methodology · Statistics 2022-03-10 Xuebin Zhao , Hong Chen , Yingjie Wang , Weifu Li , Tieliang Gong , Yulong Wang , Feng Zheng

Identifying areas where the signal is prominent is an important task in image analysis, with particular applications in brain mapping. In this work, we develop confidence regions for spatial excursion sets above and below a given level. We…

Methodology · Statistics 2025-04-18 Howon Ryu , Thomas Maullin-Sapey , Armin Schwartzman , Samuel Davenport

Testing composite null hypotheses arises in various applications, such as mediation and replicability analyses. The problem becomes more challenging in high-throughput experiments where tens of thousands of features are examined…

Methodology · Statistics 2025-04-29 Pengfei Lyu , Xianyang Zhang , Hongyuan Cao

Variable selection in cluster analysis is important yet challenging. It can be achieved by regularization methods, which realize a trade-off between the clustering accuracy and the number of selected variables by using a lasso-type penalty.…

Methodology · Statistics 2016-12-23 Marbac Matthieu , Sedki Mohammed

With the rapid growth of crowdsourcing platforms it has become easy and relatively inexpensive to collect a dataset labeled by multiple annotators in a short time. However due to the lack of control over the quality of the annotators, some…

Machine Learning · Statistics 2016-06-17 Qianqian Xu , Jiechao Xiong , Xiaochun Cao , Yuan Yao

In this paper, we have attempted to study the behaviour of the family wise error rate (FWER) for Bonferroni's procedure and false discovery rate (FDR) of the Benjamini-Hodgeberg procedure for simultaneous testing problem with equicorrelated…

Methodology · Statistics 2020-08-20 Nabaneet Das , Subir Kumar Bhandari

False discovery rate (FDR) is commonly used for correction for multiple testing in neuroimaging studies. However, when using two-tailed tests, making directional inferences about the results can lead to a vastly inflated error rate, even…

Methodology · Statistics 2025-12-16 Anderson M. Winkler , Paul A. Taylor , Thomas E. Nichols , Chris Rorden