依赖结构下高维数据协变量选择的统计方法及其在肿瘤学基因谱分类中的应用
统计理论
2019-09-13 v1 应用统计
统计方法学
统计理论
摘要
我们提出了一种在依赖结构下高维数据但观测值较少的情境中,选择与感兴趣变量相关的协变量并对其进行排序的新方法。该方法依次交织了协变量聚类、利用因子潜分析对协变量去相关、使用适应性方法的集成进行选择以及最后的排序。模拟研究显示了在各协变量聚类内部进行去相关的价值。我们首先将方法应用于37名接受化疗的晚期非小细胞肺癌患者的转录组学数据,以选择解释治疗生存结局的转录组协变量。其次,我们将方法应用于79个乳腺肿瘤样本,以定义一种新的转移生物标志物及关联基因网络的患者谱,从而实现治疗的个性化。
引用
@article{arxiv.1909.05481,
title = {A statistical methodology to select covariates in high-dimensional data under dependence. Application to the classification of genetic profiles in oncology},
author = {Bérangère Bastien and Taha Boukhobza and Hélène Dumond and Anne Gégout-Petit and Aurélie Muller-Gueudin and Charlène Thiébaut},
journal= {arXiv preprint arXiv:1909.05481},
year = {2019}
}