English
Related papers

Related papers: Sparse group variable selection for gene-environme…

200 papers

Spatial epidemiology identifies the drivers of elevated population-level disease risks, using disease counts, exposures and known confounders at the areal unit level. Poisson regression models are typically used for inference, which…

Methodology · Statistics 2026-02-03 Duncan Lee , Vinny Davies

As an important problem in causal inference, we discuss the identification and estimation of treatment effects (TEs) under limited overlap; that is, when subjects with certain features belong to a single treatment group. We use a latent…

Machine Learning · Statistics 2022-04-22 Pengzhou Wu , Kenji Fukumizu

In many longitudinal microarray studies, the gene expression levels in a random sample are observed repeatedly over time under two or more conditions. The resulting time courses are generally very short, high-dimensional, and may have…

Applications · Statistics 2013-02-26 Maurice Berk , Cheryl Hemingway , Michael Levin , Giovanni Montana

Mixed-effect models are very popular for analyzing data with a hierarchical structure, e.g. repeated observations within subjects in a longitudinal design, patients nested within centers in a multicenter design. However, recently, due to…

Methodology · Statistics 2019-05-09 Abhik Ghosh , Magne Thoresen

We consider high-dimensional regression over subgroups of observations. Our work is motivated by biomedical problems, where disease subtypes, for example, may differ with respect to underlying regression models, but sample sizes at the…

Epidemiological evidence suggests that simultaneous exposures to multiple environmental risk factors (Es) can increase disease risk larger than the additive effect of individual exposure acting alone. The interaction between a gene and…

Methodology · Statistics 2022-09-02 Shunjie Guan , Mingtao Zhao , Yuehua Cui

Unmeasured confounding presents a significant challenge in causal inference from observational studies. Classical approaches often rely on collecting proxy variables, such as instrumental variables. However, in applications where the…

Methodology · Statistics 2025-01-16 Xiaochuan Shi , Dehan Kong , Linbo Wang

We consider inference for M-estimators after model selection using a sparsity-inducing penalty. While existing methods for this task require bespoke inference procedures, we propose a simpler approach, which relies on two insights: (i)…

Methodology · Statistics 2026-01-21 Ronan Perry , Snigdha Panigrahi , Daniela Witten

We introduce a new empirical Bayes approach for large-scale multiple linear regression. Our approach combines two key ideas: (i) the use of flexible "adaptive shrinkage" priors, which approximate the nonparametric family of scale mixture of…

Methodology · Statistics 2024-06-13 Youngseok Kim , Wei Wang , Peter Carbonetto , Matthew Stephens

Motivation. Association studies have been widely used to search for associations between common genetic variants observations and a given phenotype. However, it is now generally accepted that genes and environment must be examined jointly…

Modern statistical learning algorithms are capable of amazing flexibility, but struggle with interpretability. One possible solution is sparsity: making inference such that many of the parameters are estimated as being identically 0, which…

Methodology · Statistics 2023-05-15 Nathan Wycoff , Ali Arab , Katharine M. Donato , Lisa O. Singh

Motivated by applications in neuroanatomy, we propose a novel methodology for estimating the heritability which corresponds to the proportion of phenotypic variance which can be explained by genetic factors. Estimating this quantity for…

Statistics Theory · Mathematics 2016-06-09 Anna Bonnet , Céline Lévy-Leduc , Elisabeth Gassiat , Roberto Toro , Thomas Bourgeron

We consider the problem of estimating a sparse multi-response regression function, with an application to expression quantitative trait locus (eQTL) mapping, where the goal is to discover genetic variations that influence gene-expression…

Machine Learning · Statistics 2012-10-01 Seyoung Kim , Eric P. Xing

In genetic association studies, rare variants with extremely small allele frequency play a crucial role in complex traits, and the set-based testing methods that jointly assess the effects of groups of single nucleotide polymorphisms (SNPs)…

Methodology · Statistics 2020-03-13 Shonosuke Sugasawa , Hisashi Noma

Understanding interaction effects among variables is important for regression modeling in various applications. The conventional approach of quantifying interactions as the product of variables often lacks clear interpretability, especially…

Methodology · Statistics 2026-05-21 Kexin Xie , Xinwei Deng

We present a novel approach to the formulation and the resolution of sparse Linear Discriminant Analysis (LDA). Our proposal, is based on penalized Optimal Scoring. It has an exact equivalence with penalized LDA, contrary to the multi-class…

Machine Learning · Computer Science 2012-07-03 Luis Francisco Sanchez Merchante , Yves Grandvalet , Gerrad Govaert

Motivated by genetic association studies of pleiotropy, we propose here a Bayesian latent variable approach to jointly study multiple outcomes or phenotypes. The proposed method models both continuous and binary phenotypes, and it accounts…

Applications · Statistics 2012-11-08 Lizhen Xu , Radu V. Craiu , Lei Sun

Penalized likelihood approaches are widely used for high-dimensional regression. Although many methods have been proposed and the associated theory is now well-developed, the relative efficacy of different approaches in finite-sample…

Methodology · Statistics 2020-01-29 Fan Wang , Sach Mukherjee , Sylvia Richardson , Steven M. Hill

Recent work has found that sparse autoencoders (SAEs) are an effective technique for unsupervised discovery of interpretable features in language models' (LMs) activations, by finding sparse, linear reconstructions of LM activations. We…

Machine Learning · Computer Science 2024-05-01 Senthooran Rajamanoharan , Arthur Conmy , Lewis Smith , Tom Lieberum , Vikrant Varma , János Kramár , Rohin Shah , Neel Nanda

We consider the problem of estimation of a covariance matrix for Gaussian data in a high dimensional setting. Existing approaches include maximum likelihood estimation under a pre-specified sparsity pattern, l_1-penalized loglikelihood…

Methodology · Statistics 2024-10-04 Luca Cibinel , Alberto Roverato , Veronica Vinciotti