English

Distinguishing correlation from causation using genome-wide association studies

Methodology 2018-11-28 v1 Machine Learning Machine Learning

Abstract

Genome-wide association studies (GWAS) have emerged as a rich source of genetic clues into disease biology, and they have revealed strong genetic correlations among many diseases and traits. Some of these genetic correlations may reflect causal relationships. We developed a method to quantify causal relationships between genetically correlated traits using GWAS summary association statistics. In particular, our method quantifies what part of the genetic component of trait 1 is also causal for trait 2 using mixed fourth moments E(α12α1α2)E(\alpha_1^2\alpha_1\alpha_2) and E(α22α1α2)E(\alpha_2^2\alpha_1\alpha_2) of the bivariate effect size distribution. If trait 1 is causal for trait 2, then SNPs affecting trait 1 (large α12\alpha_1^2) will have correlated effects on trait 2 (large α1α2\alpha_1\alpha_2), but not vice versa. We validated this approach in extensive simulations. Across 52 traits (average N=331N=331k), we identified 30 putative genetically causal relationships, many novel, including an effect of LDL cholesterol on decreased bone mineral density. More broadly, we demonstrate that it is possible to distinguish between genetic correlation and causation using genetic association data.

Keywords

Cite

@article{arxiv.1811.08803,
  title  = {Distinguishing correlation from causation using genome-wide association studies},
  author = {Luke J. O'Connor and Alkes L. Price},
  journal= {arXiv preprint arXiv:1811.08803},
  year   = {2018}
}

Comments

Machine Learning for Health (ML4H) Workshop at NeurIPS 2018 arXiv:1811.07216

R2 v1 2026-06-23T05:23:36.715Z