中文

利用 cat 分数和错误非发现率控制进行组学预测问题中的特征选择

应用统计 2010-10-11 v4 机器学习

摘要

我们重新审视线性判别分析 (LDA) 中的特征选择问题,即特征相关的情况。首先,我们引入了多类 LDA 预测函数的合并质心公式,其中马氏变换预测变量的相对权重由相关性调整的 t 分数 (cat 分数) 给出。其次,对于特征选择,我们建议通过控制错误非发现率 (FNDR) 来阈值化 cat 分数。第三,分类器的训练基于相关性和方差的 James-Stein 收缩估计,其中正则化参数通过解析选择而无需重采样。总体而言,这为具有自然特征选择的高维预测提供了一个有效且计算成本低廉的框架。所提出的收缩判别程序已在 R 包 "sda" 中实现,可从 R 仓库 CRAN 获取。

关键词

引用

@article{arxiv.0903.2003,
  title  = {Feature selection in omics prediction problems using cat scores and false nondiscovery rate control},
  author = {Miika Ahdesmäki and Korbinian Strimmer},
  journal= {arXiv preprint arXiv:0903.2003},
  year   = {2010}
}

备注

Published in at http://dx.doi.org/10.1214/09-AOAS277 the Annals of Applied Statistics (http://www.imstat.org/aoas/) by the Institute of Mathematical Statistics (http://www.imstat.org)