English
Related papers

Related papers: Disentangling the effects of traits with shared cl…

200 papers

Motivated by genome-wide association studies, we consider a standard linear model with one additional random effect in situations where many predictors have been collected on the same subjects and each predictor is analyzed separately.…

Applications · Statistics 2013-04-24 Matti Pirinen , Peter Donnelly , Chris C. A. Spencer

Cluster-randomized trials (CRTs) are widely used to evaluate interventions delivered at the clinic, practice, or community level. Although standard analyses typically target average treatment effects, such summaries mask potentially…

Methodology · Statistics 2026-04-16 Changjun Li , Xi Fang , Michael O. Harhay , Andrew B. Forbes , F. Perry Wilson , Guangyu Tong , Fan Li

It is widely recognized nowadays that complex diseases are caused by, amongst the others, multiple genetic factors. The recent advent of genome-wide association study (GWA) has triggered a wave of research aimed at discovering genetic…

Applications · Statistics 2011-09-30 Nguyen Xuan Vinh

When drawing causal inferences about the effects of multiple treatments on clustered survival outcomes using observational data, we need to address implications of the multilevel data structure, multiple treatments, censoring and unmeasured…

Methodology · Statistics 2022-02-18 Liangyuan Hu , Jiayi Ji , Ronald D. Ennis , Joseph W. Hogan

Risk prediction models using genetic data have seen increasing traction in genomics. However, most of the polygenic risk models were developed using data from participants with similar (mostly European) ancestry. This can lead to biases in…

Machine Learning · Computer Science 2022-05-11 Prashnna K Gyawali , Yann Le Guen , Xiaoxia Liu , Hua Tang , James Zou , Zihuai He

In many applications, data cluster. Failing to take the cluster structure into consideration generally leads to underestimated variances of point estimators and inflated type I errors in hypothesis tests. Many circumstance-dependent…

Methodology · Statistics 2025-07-21 Jiahua Chen , Pengfei Li , Yukun Liu , James V. Zidek

Mendelian randomization (MR) is an epidemiological method that can be used to strengthen causal inference regarding the relationship between a modifiable environmental exposure and a medically relevant trait and to estimate the magnitude of…

Quantitative Methods · Quantitative Biology 2023-08-30 David M Evans , George Davey Smith , Gunn-Helen Moen

In the framework of model-based clustering, a model allowing several latent class variables is proposed. This model assumes that the distribution of the observed data can be factorized into several independent blocks of variables. Each…

Methodology · Statistics 2018-01-23 Matthieu Marbac , Vincent Vandewalle

Complex diseases are multifactorial traits caused by both genetic and environmental factors. They represent the most part of human diseases and include those with largest prevalence and mortality (cancer, heart disease, obesity, etc.).…

Quantitative Methods · Quantitative Biology 2009-05-08 Roberto Amato , Michele Pinelli , Daniel D'Andrea , Gennaro Miele , Mario Nicodemi , Giancarlo Raiconi , Sergio Cocozza

In univariate data, there exist standard procedures for identifying dominating features that produce the largest observations. However, in the multivariate setting, the situation is quite different. This paper aims to provide tools and…

Statistics Theory · Mathematics 2023-08-29 Miriam Hägele , Jaakko Lehtomaa

Cluster analysis methods seek to partition a data set into homogeneous subgroups. It is useful in a wide variety of applications, including document processing and modern genetics. Conventional clustering methods are unsupervised, meaning…

Methodology · Statistics 2014-07-11 Eric Bair

To date, efforts to produce high-quality polygenic risk scores from genome-wide studies of common disease have focused on estimating and aggregating the effects of multiple SNPs. Here we propose a novel statistical approach for genetic risk…

Quantitative Methods · Quantitative Biology 2014-05-13 David Golan , Saharon Rosset

Mendelian Randomization (MR) is a prominent observational epidemiological research method designed to address unobserved confounding when estimating causal effects. However, core assumptions -- particularly the independence between…

Machine Learning · Computer Science 2026-02-24 Shimeng Huang , Matthew Robinson , Francesco Locatello

A theoretical framework is presented for a (copula-based) notion of dissimilarity between continuous random vectors and its main properties are studied. The proposed dissimilarity assigns the smallest value to a pair of random vectors that…

Methodology · Statistics 2021-02-04 Sebastian Fuchs , F. Marta L. Di Lascio , Fabrizio Durante

Telling apart the cause and effect between two random variables with purely observational data is a challenging problem that finds applications in various scientific disciplines. A key principle utilized in this task is the algorithmic…

Machine Learning · Computer Science 2025-08-15 Quang-Duy Tran , Bao Duong , Phuoc Nguyen , Thin Nguyen

In clustering we normally output one cluster variable for each datapoint. However it is not necessarily the case that there is only one way to partition a given dataset into cluster components. For example, one could cluster objects by…

Machine Learning · Computer Science 2019-12-05 Matthew Willetts , Stephen Roberts , Chris Holmes

Mendelian randomization (MR) has become a popular approach to study causal effects by using genetic variants as instrumental variables. We propose a new MR method, GENIUS-MAWII, which simultaneously addresses the two salient phenomena that…

Methodology · Statistics 2024-02-27 Ting Ye , Zhonghua Liu , Baoluo Sun , Eric Tchetgen Tchetgen

We address estimation of intervention effects in experimental designs in which (a) interventions are assigned at the cluster level; (b) clusters are selected to form pairs, matched on observed characteristics; and (c) intervention is…

Methodology · Statistics 2014-11-24 Zhenke Wu , Constantine E. Frangakis , Thomas A. Louis , Daniel O. Scharfstein

The results from Genome-Wide Association Studies (GWAS) on thousands of phenotypes provide an unprecedented opportunity to infer the causal effect of one phenotype (exposure) on another (outcome). Mendelian randomization (MR), an…

Methodology · Statistics 2019-04-30 Jia Zhao , Jingsi Ming , Xianghong Hu , Gang Chen , Jin Liu , Can Yang

Batch effects represent a major confounder in genomic diagnostics. In copy number variant (CNV) detection from NGS, many algorithms compare read depth between test samples and a reference sample, assuming they are process-matched. When this…

Genomics · Quantitative Biology 2026-01-16 Austin Talbot , Yue Ke