中文
相关论文

相关论文: A Powerful and Precise Feature-level Filter using …

200 篇论文

Power and reproducibility are key to enabling refined scientific discoveries in contemporary big data applications with general high-dimensional nonlinear models. In this paper, we provide theoretical foundations on the power and robustness…

统计理论 · 数学 2017-09-04 Yingying Fan , Emre Demirkaya , Gaorong Li , Jinchi Lv

Genomics biobanks are information treasure troves with thousands of phenotypes (e.g., diseases, traits) and millions of single nucleotide polymorphisms (SNPs). The development of methodologies that provide reproducible discoveries is…

统计方法学 · 统计学 2024-10-08 Jasin Machkour , Michael Muma , Daniel P. Palomar

Deep neural networks (DNN) have been used successfully in many scientific problems for their high prediction accuracy, but their application to genetic studies remains challenging due to their poor interpretability. In this paper, we…

机器学习 · 计算机科学 2021-10-01 Peyman H. Kassani , Fred Lu , Yann Le Guen , Zihuai He

In many research fields, researchers aim to identify significant associations between a set of explanatory variables and a response while controlling the FDR. The Knockoff filter has been recently proposed in the frequentist paradigm to…

统计方法学 · 统计学 2026-04-22 Lorenzo Focardi-Olmi , Anna Gottard , Michele Guindani , Marina Vannucci

We propose a new method to learn the structure of a Gaussian graphical model with finite sample false discovery rate control. Our method builds on the knockoff framework of Barber and Cand\`{e}s for linear models. We extend their approach…

统计方法学 · 统计学 2021-04-20 Jinzhou Li , Marloes H. Maathuis

We develop a new class of distribution--free multiple testing rules for false discovery rate (FDR) control under general dependence. A key element in our proposal is a symmetrized data aggregation (SDA) approach to incorporating the…

统计方法学 · 统计学 2021-05-27 Lilun Du , Xu Guo , Wenguang Sun , Changliang Zou

Early and accurate detection of Alzheimer's disease (AD) remains a major challenge in medical diagnosis due to its subtle onset and progressive nature. This research introduces an explainable ensemble learning Framework designed to classify…

机器学习 · 计算机科学 2026-03-06 Nishan Mitra

In many real world problems, features do not act alone but in combination with each other. For example, in genomics, diseases might not be caused by any single mutation but require the presence of multiple mutations. Prior work on feature…

机器学习 · 计算机科学 2023-01-12 Fergus Imrie , Alexander Norcliffe , Pietro Lio , Mihaela van der Schaar

We develop a flexible feature selection framework based on deep neural networks that approximately controls the false discovery rate (FDR), a measure of Type-I error. The method applies to architectures whose first layer is fully connected.…

机器学习 · 统计学 2026-02-10 Kazuma Sawaya

The knockoff filter of Barber and Candes (arXiv:1404.5609) is a flexible framework for multiple testing in supervised learning models, based on introducing synthetic predictor variables to control the false discovery rate (FDR). Using the…

统计方法学 · 统计学 2024-11-26 Yixiang Luo , William Fithian , Lihua Lei

The rapid generation of complex, highly skewed, and zero-inflated multi-source count data poses significant challenges for variable selection, particularly in biomedical domains like tumor development and metabolic dysregulation. To address…

应用统计 · 统计学 2025-11-11 Shan Tang , Shanjun Mao , Shourong Ma , Falong Tan

Many contemporary large-scale applications involve building interpretable models linking a large set of potential covariates to a response in a nonlinear fashion, such as when the response is binary. Although this modeling problem has been…

统计方法学 · 统计学 2017-12-13 Emmanuel Candes , Yingying Fan , Lucas Janson , Jinchi Lv

Automatic identification and categorization of Alzheimer's patients and the ability to distinguish between different levels of this disease can be very helpful to the research community in this field, since other non-automatic approaches…

信号处理 · 电气工程与系统科学 2019-04-17 Esmaeil Seraj , Mehran Yazdi , Nastaran Shahparian

This paper introduces a machine for sampling approximate model-X knockoffs for arbitrary and unspecified data distributions using deep generative models. The main idea is to iteratively refine a knockoff sampling mechanism until a criterion…

统计方法学 · 统计学 2020-03-03 Yaniv Romano , Matteo Sesia , Emmanuel J. Candès

Controlling the false discovery rate (FDR) in high-dimensional variable selection requires balancing rigorous error control with statistical power. Existing methods with provable guarantees are often overly conservative, creating a…

统计方法学 · 统计学 2026-02-06 Arnau Vilella , Jasin Machkour , Michael Muma , Daniel P. Palomar

Genomic data are subject to various sources of confounding, such as demographic variables, biological heterogeneity, and batch effects. To identify genomic features associated with a variable of interest in the presence of confounders, the…

统计方法学 · 统计学 2025-12-08 Asmita Roy , Jun Chen , Xianyang Zhang

Conditional testing via the knockoff framework allows one to identify -- among large number of possible explanatory variables -- those that carry unique information about an outcome of interest, and also provides a false discovery rate…

统计方法学 · 统计学 2024-03-05 Benjamin B Chu , Jiaqi Gu , Zhaomeng Chen , Tim Morrison , Emmanuel Candes , Zihuai He , Chiara Sabatti

Recommendation systems and computing advertisements have gradually entered the field of academic research from the field of commercial applications. Click-through rate prediction is one of the core research issues because the prediction…

机器学习 · 计算机科学 2019-02-26 Li Zhang , Weichen Shen , Shijian Li , Gang Pan

Alzheimer disease is a severe brain disorder that causes harm in various brain areas and leads to memory damage. The limited availability of labeled medical data poses a significant challenge for accurate Alzheimer disease detection. There…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Safa Ben Atitallah , Maha Driss , Wadii Boulila , Anis Koubaa

The current state-of-the-art deep neural networks (DNNs) for Alzheimer's Disease diagnosis use different biomarker combinations to classify patients, but do not allow extracting knowledge about the interactions of biomarkers. However, to…

机器学习 · 计算机科学 2021-09-28 Raphael Ronge , Kwangsik Nho , Christian Wachinger , Sebastian Pölsterl