English
Related papers

Related papers: Aggregating Knockoffs for False Discovery Rate Con…

200 papers

Filters (such as Bloom Filters) are data structures that speed up network routing and measurement operations by storing a compressed representation of a set. Filters are space efficient, but can make bounded one-sided errors: with tunable…

Data Structures and Algorithms · Computer Science 2021-05-25 Tsvi Kopelowitz , Samuel McCauley , Ely Porat

Understanding effect modification -- how treatment effects vary across subpopulations -- is practically important in observational studies, as it helps identify which subgroups are likely to benefit from a given treatment. In this paper, we…

Methodology · Statistics 2026-05-12 Yu Gui , Dylan S Small , Zhimei Ren

The traditional framework for feature selection treats all features as costing the same amount. However, in reality, a scientist often has considerable discretion regarding which variables to measure, and the decision involves a tradeoff…

Methodology · Statistics 2023-02-14 Guo Yu , Daniela Witten , Jacob Bien

The analysis of environmental mixtures is of growing importance in environmental epidemiology, and one of the key goals in such analyses is to identify exposures and their interactions that are associated with adverse health outcomes.…

Methodology · Statistics 2021-03-22 Srijata Samanta , Joseph Antonelli

This paper presents a survey on some recent advances for the type I error rate control in multiple testing methodology. We consider the problem of controlling the $k$-family-wise error rate (kFWER, probability to make $k$ false discoveries…

Methodology · Statistics 2011-03-15 Etienne Roquain

Node classification is a substantial problem in graph-based fraud detection. Many existing works adopt Graph Neural Networks (GNNs) to enhance fraud detectors. While promising, currently most GNN-based fraud detectors fail to generalize to…

Artificial Intelligence · Computer Science 2023-02-22 Yuchen Wang , Jinghui Zhang , Zhengjie Huang , Weibin Li , Shikun Feng , Ziheng Ma , Yu Sun , Dianhai Yu , Fang Dong , Jiahui Jin , Beilun Wang , Junzhou Luo

In this era of "big" data, not only the large amount of data keeps motivating distributed computing, but concerns on data privacy also put forward the emphasis on distributed learning. To conduct feature selection and to control the false…

Methodology · Statistics 2020-08-11 Yu Gui

Understanding food acquisition is crucial for developing strategies to combat food insecurity, a major public health concern. The emergence of large-scale mobile location data (typically exemplified by GPS data), which captures people's…

Computers and Society · Computer Science 2025-08-22 Duanya Lyu , Luyu Liu , Catherine Campbell , Yuxuan Zhang , Xiang Yan

Atrophic gastritis is a significant risk factor for developing gastric cancer. The incorporation of machine learning algorithms can efficiently elevate the possibility of accurately detecting atrophic gastritis. Nevertheless, when the…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Zinan Xiong , Shuijiao Chen , Yizhe Zhang , Yu Cao , Benyuan Liu , Xiaowei Liu

The popularity of penalized regression in high-dimensional data analysis has led to a demand for new inferential tools for these models. False discovery rate control is widely used in high-dimensional hypothesis testing, but has only…

Methodology · Statistics 2019-01-24 Ryan Miller , Patrick Breheny

We propose a novel multiple testing methodology for controlling the false discovery rate (FDR) in high-dimensional linear models that integrates model-X knockoff techniques with debiased penalized regression estimators. At the foundation of…

Methodology · Statistics 2026-03-17 Jinyuan Chang , Chenlong Li , Cheng Yong Tang , Zhengtian Zhu

Deep-learning-based pipelines have shown the potential to revolutionalize microscopy image diagnostics by providing visual augmentations to a trained pathology expert. However, to match human performance, the methods rely on the…

The recently proposed fixed-X knockoff is a powerful variable selection procedure that controls the false discovery rate (FDR) in any finite-sample setting, yet its theoretical insights are difficult to show beyond Gaussian linear models.…

Methodology · Statistics 2023-11-28 Han Su , Panxu Yuan , Qingyang Sun , Mengxi Yi , Gaorong Li

Crohn's disease (CD) is associated with an ecological imbalance of the intestinal microbiota, consisting of hundreds of species. The underlying complexity as well as individual differences between patients contributes to the difficulty to…

Molecular Networks · Quantitative Biology 2017-09-19 Eugen Bauer , Ines Thiele

Consider the problem of simultaneously testing null hypotheses H_1,...,H_s. The usual approach to dealing with the multiplicity problem is to restrict attention to procedures that control the familywise error rate (FWER), the probability of…

Statistics Theory · Mathematics 2007-06-13 E. L. Lehmann , Joseph P. Romano

In multiple testing several criteria to control for type I errors exist. The false discovery rate, which evaluates the expected proportion of false discoveries among the rejected null hypotheses, has become the standard approach in this…

Methodology · Statistics 2023-11-03 Jacobo de Uña-Álvarez

Controlling false discovery rate (FDR) while leveraging the side information of multiple hypothesis testing is an emerging research topic in modern data science. Existing methods rely on the test-level covariates while ignoring possible…

Machine Learning · Statistics 2021-01-26 Lin Qiu , Nils Murrugarra-Llerena , Vítor Silva , Lin Lin , Vernon M. Chinchilli

Discrete data such as counts of microbiome taxa resulting from next-generation sequencing are routinely encountered in bioinformatics. Taxa count data in microbiome studies are typically high-dimensional, over-dispersed, and can only reveal…

Methodology · Statistics 2022-06-23 Yuan Fang , Sanjeena Subedi

Stability and reproducibility are essential considerations in various applications of statistical methods. False Discovery Rate (FDR) control methods are able to control false signals in scientific discoveries. However, many FDR control…

Methodology · Statistics 2025-12-22 Jiajun Sun , Zhanrui Cai , Wei Zhong

Data augmentation plays a key role in modern machine learning pipelines. While numerous augmentation strategies have been studied in the context of computer vision and natural language processing, less is known for other data modalities.…

Machine Learning · Statistics 2022-05-23 Elliott Gordon-Rodriguez , Thomas P. Quinn , John P. Cunningham