缺失数据关联研究中的同时SNP识别
应用统计
2012-07-04 v2
摘要
关联检验旨在发现基因型(通常是单核苷酸多态性,即SNP)与表型(属性或性状)之间的潜在关系。关联检验中使用的典型大数据集通常包含缺失值。标准统计方法要么使用相对简单的假设对缺失值进行插补,要么删除它们,或两者兼用,这可能会产生有偏结果。这里我们描述了贝叶斯分层模型BAMD(缺失数据贝叶斯关联)。BAMD是一个吉布斯采样器,其中基于数据集中所有可用信息对缺失值进行多重插补。我们估计参数并证明每次迭代更新一个SNP保持了马尔可夫链的遍历性质,同时提高了计算速度。我们还在BAMD中实现了一个模型选择选项,能够潜在检测SNP相互作用。模拟表明,在基因型数据缺失的情况下,SNP效应的无偏估计得以恢复。此外,我们验证了先前使用基于家系的方法报道的SNP与碳同位素辨别表型之间的关联,并发现了一个与该性状相关的额外SNP。BAMD作为R包可从http://cran.r-project.org/package=BAMD获取。
引用
@article{arxiv.1207.0280,
title = {Simultaneous SNP identification in association studies with missing data},
author = {Zhen Li and Vikneswaran Gopal and Xiaobo Li and John M. Davis and George Casella},
journal= {arXiv preprint arXiv:1207.0280},
year = {2012}
}
备注
Published in at http://dx.doi.org/10.1214/11-AOAS516 the Annals of Applied Statistics (http://www.imstat.org/aoas/) by the Institute of Mathematical Statistics (http://www.imstat.org)