English
Related papers

Related papers: StaRQR-K: False Discovery Rate Controlled Regional…

200 papers

Genomics biobanks are information treasure troves with thousands of phenotypes (e.g., diseases, traits) and millions of single nucleotide polymorphisms (SNPs). The development of methodologies that provide reproducible discoveries is…

Methodology · Statistics 2024-10-08 Jasin Machkour , Michael Muma , Daniel P. Palomar

Unsupervised learning on high-dimensional RNA-seq data can reveal molecular subtypes beyond standard labels. We combine an autoencoder-based representation with clustering and stability analysis to search for rare but reproducible genomic…

Machine Learning · Computer Science 2025-11-18 Alaa Mezghiche

Controlling the False Discovery Rate (FDR) is critical for reproducible variable selection, especially given the prevalence of complex predictive modeling. The recent Split Knockoff method, an extension of the canonical Knockoffs framework,…

Methodology · Statistics 2025-09-05 Yang Cao , Hangyu Lin , Xinwei Sun , Yuan Yao

Controlling the False Discovery Rate (FDR) in a variable selection procedure is critical for reproducible discoveries, and it has been extensively studied in sparse linear models. However, it remains largely open in scenarios where the…

Methodology · Statistics 2023-11-16 Yang Cao , Xinwei Sun , Yuan Yao

Variable selection has been widely used in data analysis for the past decades, and it becomes increasingly important in the Big Data era as there are usually hundreds of variables available in a dataset. To enhance interpretability of a…

Methodology · Statistics 2020-08-17 Yuxiang Xie , Kwun Chuen Gary Chan

A new statistical procedure (Model-X \cite{candes2018}) has provided a way to identify important factors using any supervised learning method controlling for FDR. This line of research has shown great potential to expand the horizon of…

Methodology · Statistics 2018-10-01 Ying Liu , Cheng Zheng

Colorectal cancer remains a major global health concern, with early detection being pivotal for improving patient outcomes. In this study, we leveraged high throughput methylation profiling of cellfree DNA to identify and validate…

Genomics · Quantitative Biology 2025-05-19 Kartavya Mathur , Shipra Jain , Nisha Bajiya , Nishant Kumar , Gajendra P. S. Raghava

In many fields of science, we observe a response variable together with a large number of potential explanatory variables, and would like to be able to discover which variables are truly associated with the response. At the same time, we…

Methodology · Statistics 2015-10-15 Rina Foygel Barber , Emmanuel J. Candès

Predicting drug responses using genetic and transcriptomic features is crucial for enhancing personalized medicine. In this study, we implemented an ensemble of machine learning algorithms to analyze the correlation between genetic and…

Genomics · Quantitative Biology 2025-07-04 Johannes Schlüter , Alexander Schönhuth

We propose a prediction procedure for the functional linear quantile regression model by using partial quantile covariance techniques and develop a simple partial quantile regression (SIMPQR) algorithm to efficiently extract partial…

Methodology · Statistics 2015-11-03 Dengdeng Yu , Linglong Kong , Ivan Mizera

We propose the group knockoff filter, a method for false discovery rate control in a linear regression setting where the features are grouped, and we would like to select a set of relevant groups which have a nonzero effect on the response.…

Methodology · Statistics 2016-02-12 Ran Dai , Rina Foygel Barber

Multivariate statistics are often available as well as necessary in hypothesis tests. We study how to use such statistics to control not only false discovery rate (FDR) but also positive FDR (pFDR) with good power. We show that FDR can be…

Statistics Theory · Mathematics 2008-05-21 Zhiyi Chi

In genome-wide association studies, hundreds of thousands of genetic features (genes, proteins, etc.) in a given case-control population are tested to verify existence of an association between each genetic marker and a specific disease. A…

Applications · Statistics 2021-02-01 Ali Karimnezhad

We describe a series of algorithms that efficiently implement Gaussian model-X knockoffs to control the false discovery rate on large scale feature selection problems. Identifying the knockoff distribution requires solving a large scale…

Machine Learning · Computer Science 2020-06-17 Armin Askari , Quentin Rebjock , Alexandre d'Aspremont , Laurent El Ghaoui

DNA methylation is a well-studied genetic modification that regulates gene transcription of Eukaryotes. Its alternations have been recognized as a significant component of cancer development. In this study, we use the DNA methylation 450k…

Tissues and Organs · Quantitative Biology 2021-01-05 Shen Jia , Yulin Zhang , Yiming Mao , Jiawei Gao , Yixuan Chen , Yuxuan Jiang , Haochen Luo , Kebo Lv , Jionglong Su

Cancer development is associated with aberrant DNA methylation, including increased stochastic variability. Statistical tests for discovering cancer methylation biomarkers have focused on changes in mean methylation. To improve the power of…

Methodology · Statistics 2023-06-27 James Y. Dai , Heng Chen , Xiaoyu Wang , Wei Sun , Ying Huang , William M. Grady , Ziding Feng

Signal identification in large-dimensional settings is a challenging problem in biostatistics. Recently, the method of higher criticism (HC) was shown to be an effective means for determining appropriate decision thresholds. Here, we study…

Methodology · Statistics 2012-12-21 Bernd Klaus , Korbinian Strimmer

Thanks to its fine balance between model flexibility and interpretability, the nonparametric additive model has been widely used, and variable selection for this type of model has been frequently studied. However, none of the existing…

Methodology · Statistics 2022-01-10 Xiaowu Dai , Xiang Lyu , Lexin Li

Research on the localization of the genetic basis associated with diseases or traits has been widely conducted in the last a few decades. Scan methods have been developed for region-based analysis in whole-genome association studies,…

Methodology · Statistics 2024-10-31 Wei Zhang , Fan Wang , Fang Yao

E-values have been the dominant statistic for protein sequence analysis for the past two decades: from identifying statistically significant local sequence alignments to evaluating matches to hidden Markov models describing protein domain…

Genomics · Quantitative Biology 2016-02-17 Alejandro Ochoa , John D. Storey , Manuel Llinás , Mona Singh