English
Related papers

Related papers: Second-order group knockoffs with applications to …

200 papers

Drug development is a very costly and lengthy process, while repositioned or repurposed drugs could be brought into clinical practice within a shorter time-frame and at a much reduced cost. The past decade has observed a massive growth in…

Genomics · Quantitative Biology 2019-11-14 Alexandria Lau , Hon-Cheong So

Multivariate Gaussian distributions enjoy Gaussian conditional distributions that makes conditioning easy: conditioning boils down to implementing analytical formulae for conditional means and covariances. For more general distributions,…

Methodology · Statistics 2026-03-26 Antoine Faul , David Ginsbourger , Ben Spycher

Canonical correlation analysis (CCA) is a classic statistical method for discovering latent co-variation that underpins two or more observed random vectors. Several extensions and variations of CCA have been proposed that have strengthened…

Machine Learning · Computer Science 2023-12-22 Paris A. Karakasis , Nicholas D. Sidiropoulos

Vine copulas are a flexible tool for high-dimensional dependence modeling. In this article, we discuss the generation of approximate model-X knockoffs with vine copulas. It is shown how Gaussian knockoffs can be generalized to Gaussian…

Methodology · Statistics 2022-10-21 Malte S. Kurz

In many applications, we need to study a linear regression model that consists of a response variable and a large number of potential explanatory variables and determine which variables are truly associated with the response. In 2015,…

Methodology · Statistics 2019-07-23 Jiajie Chen , Anthony Hou , Thomas Y. Hou

Disease-gene association through Genome-wide association study (GWAS) is an arduous task for researchers. Investigating single nucleotide polymorphisms (SNPs) that correlate with specific diseases needs statistical analysis of associations.…

Quantitative Methods · Quantitative Biology 2020-12-21 Sezin Kircali Ata , Min Wu , Yuan Fang , Le Ou-Yang , Chee Keong Kwoh , Xiao-Li Li

Variable selection properties of procedures utilizing penalized-likelihood estimates is a central topic in the study of high dimensional linear regression problems. Existing literature emphasizes the quality of ranking of the variables by…

Statistics Theory · Mathematics 2022-04-28 Asaf Weinstein , Weijie J. Su , Małgorzata Bogdan , Rina F. Barber , Emmanuel J. Candès

Genome-wide association studies (GWAS) have identified hundreds of loci at very stringent levels of statistical significance across many different human traits. However, it is now clear that very large samples (n~10^4-10^5) are needed to…

Genomics · Quantitative Biology 2013-08-20 Inti Pedroso

There is a growing need for unbiased clustering methods, ideally automated. We have developed a topology-based analysis tool called Two-Tier Mapper (TTMap) to detect subgroups in global gene expression datasets and identify their…

Genomics · Quantitative Biology 2018-01-08 Rachel Jeitziner , Mathieu Carrière , Jacques Rougemont , Steve Oudot , Kathryn Hess , Cathrin Brisken

In a group testing scheme, a set of tests is designed to identify a small number $t$ of defective items that are present among a large number $N$ of items. Each test takes as input a group of items and produces a binary output indicating…

Information Theory · Computer Science 2016-11-17 Arya Mazumdar

The study in group testing aims to develop strategies to identify a small set of defective items among a large population using a few pooled tests. The established techniques have been highly beneficial in a broad spectrum of applications…

Information Theory · Computer Science 2025-01-23 Venkata Gandikota , Nikita Polyanskii , Haodong Yang

In clinical trials, there is potential to improve precision and reduce the required sample size by appropriately adjusting for baseline variables in the statistical analysis. This is called covariate adjustment. Despite recommendations by…

Methodology · Statistics 2022-06-20 Kelly Van Lancker , Joshua Betz , Michael Rosenblum

As genomic research has grown increasingly popular in recent years, dataset sharing has remained limited due to privacy concerns. This limitation hinders the reproducibility and validation of research outcomes, both of which are essential…

Cryptography and Security · Computer Science 2025-04-02 Yuzhou Jiang , Tianxi Ji , Erman Ayday

We propose a unified theoretical framework for studying the robustness of the model-X knockoffs framework by investigating the asymptotic false discovery rate (FDR) control of the practically implemented approximate knockoffs procedure.…

Machine Learning · Statistics 2025-02-11 Yingying Fan , Lan Gao , Jinchi Lv , Xiaocong Xu

Let X; Z be r and s-dimensional covariates, respectively, used to model the response variable Y as Y = m(X;Z) + \sigma(X;Z)\epsilon. We develop an ANOVA-type test for the null hypothesis that Z has no influence on the regression function,…

Methodology · Statistics 2016-11-11 Adriano Zanin Zambom , Michael G. Akritas

A general framework for dealing with both linear regression and clustering problems is described. It includes Gaussian clusterwise linear regression analysis with random covariates and cluster analysis via Gaussian mixture models with…

Methodology · Statistics 2015-10-13 Giuliano Galimberti , Annamaria Manisi , Gabriele Soffritti

2 Diabetes is a leading worldwide public health concern, and its increasing prevalence has significant health and economic importance in all nations. The condition is a multifactorial disorder with a complex aetiology. The genetic…

Machine Learning · Computer Science 2018-08-30 Basma Abdulaimma , Paul Fergus , Carl Chalmers

The protection of privacy of individual-level information in genome-wide association study (GWAS) databases has been a major concern of researchers following the publication of "an attack" on GWAS data by Homer et al. (2008) Traditional…

Applications · Statistics 2014-02-10 Fei Yu , Stephen E. Fienberg , Aleksandra Slavković , Caroline Uhler

Standard high-dimensional factor models assume that the comovements in a large set of variables could be modeled using a small number of latent factors that affect all variables. In many relevant applications in economics and finance,…

Econometrics · Economics 2022-02-08 Antoine Djogbenou , Razvan Sufana

We propose a new method to learn the structure of a Gaussian graphical model with finite sample false discovery rate control. Our method builds on the knockoff framework of Barber and Cand\`{e}s for linear models. We extend their approach…

Methodology · Statistics 2021-04-20 Jinzhou Li , Marloes H. Maathuis