中文

高维线性回归中Lasso选择的稀疏性与偏差

统计理论 2008-08-08 v1 统计理论

摘要

Meinshausen和Buhlmann [Ann. Statist. 34 (2006) 1436–1462] 表明,对于高斯图模型中的邻域选择,在邻域稳定性条件下,即使变量数量远大于样本量,LASSO也是一致的。Zhao和Yu [(2006) J. Machine Learning Research 7 2541–2567] 将线性回归背景下的邻域稳定性条件形式化为强不可表示条件。该文表明,在此条件下,只要非零回归系数以一定速率远离零,LASSO就能精确选择出非零回归系数的集合。在本文中,理想模型之外的回归系数被假定为很小,但不一定为零。在设计变量相关性的稀疏Riesz条件下,我们证明了LASSO选择了正确维数阶的模型,将所选模型的偏差控制在由小回归系数贡献和阈值偏差决定的水平,并选择了所有阶数大于所选模型偏差的系数。此外,作为LASSO在模型选择中这种速率一致性的结果,我们证明了在给定条件下,均值响应的误差平方和以及回归系数的α\ell_{\alpha}损失以最佳可能速率收敛。我们结果的一个有趣方面是,对于某些随机依赖设计,变量数量的对数可以与样本量同阶。

关键词

引用

@article{arxiv.0808.0967,
  title  = {The sparsity and bias of the Lasso selection in high-dimensional linear regression},
  author = {Cun-Hui Zhang and Jian Huang},
  journal= {arXiv preprint arXiv:0808.0967},
  year   = {2008}
}

备注

Published in at http://dx.doi.org/10.1214/07-AOS520 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)