定制化训练及其在癌症组织质谱成像中的应用
应用统计
2016-02-01 v1 定量方法
摘要
我们提出了一种简单且可解释的预测策略,用于在模型拟合时测试数据的特征可用的情况下对测试数据进行预测。我们的提议——定制化训练(customized training)——对数据进行聚类,以找到靠近每个测试点的训练点,然后在每个训练簇中分别拟合一个 -正则化模型(lasso)。该方法结合了 近邻的局部适应性与 lasso 的可解释性。尽管我们使用 lasso 进行模型拟合,但任何监督学习方法都可应用于定制化训练集。我们将该方法应用于一个来自胃癌检测持续合作项目的质谱成像数据集,这展示了该技术的效力与可解释性。我们的想法简单,但在数据具有某些底层结构的情况下可能有用。
引用
@article{arxiv.1601.07994,
title = {Customized training with an application to mass spectrometric imaging of cancer tissue},
author = {Scott Powers and Trevor Hastie and Robert Tibshirani},
journal= {arXiv preprint arXiv:1601.07994},
year = {2016}
}
备注
Published at http://dx.doi.org/10.1214/15-AOAS866 in the Annals of Applied Statistics (http://www.imstat.org/aoas/) by the Institute of Mathematical Statistics (http://www.imstat.org)