中文

在基于质谱的无标记定量蛋白质组学差异分析中考虑多重插补引起的变异性

统计方法学 2022-09-08 v1 定量方法 应用统计

摘要

在无标记定量蛋白质组学中,插补缺失值是常见做法。插补旨在用用户定义的值替换缺失值。然而,插补过程本身在插补的下游可能未被充分考量,因为插补后的数据集常被视为原本就是完整的。因此,由插补引起的不确定性未被恰当计入。我们提供了一种严格的多重插补策略,借助 Rubin 规则实现对参数变异性偏差更小的估计。随后利用贝叶斯层次模型对基于插补的肽段强度方差估计量进行调制。该估计量最终被纳入调制 t 检验统计量中以给出差异分析结果。该流程在定量数据集的肽段和蛋白质水平均可使用。对于基于肽段水平定量数据的蛋白质水平结果,还包含了一个聚合步骤。我们的方法名为 mi4p,在模拟和真实数据集上均与 DAPAR R 包中实现的当前最优 limma 流程进行了比较。我们观察到灵敏度与特异性之间存在权衡,而 mi4p 在 F 值方面的整体表现优于 DAPAR。

关键词

引用

@article{arxiv.2108.07086,
  title  = {Accounting for multiple imputation-induced variability for differential analysis in mass spectrometry-based label-free quantitative proteomics},
  author = {Marie Chion and Christine Carapito and Frédéric Bertrand},
  journal= {arXiv preprint arXiv:2108.07086},
  year   = {2022}
}

备注

The methodology here described is implemented under the R environment and can be found on GitHub: https://github.com/mariechion/mi4p. The R scripts which led to the results presented here can also be found on this repository. The real datasets are available on ProteomeXchange under the dataset identifiers PXD003841 and PXD027800