A Risk Ratio Comparison of $l_0$ and $l_1$ Penalized Regression
Abstract
There has been an explosion of interest in using -regularization in place of -regularization for feature selection. We present theoretical results showing that while -penalized linear regression never outperforms -regularization by more than a constant factor, in some cases using an penalty is infinitely worse than using an penalty. We also show that the "optimal" solutions are often inferior to solutions found using stepwise regression. We also compare algorithms for solving these two problems and show that although solutions can be found efficiently for the problem, the "optimal" solutions are often inferior to solutions found using greedy classic stepwise regression. Furthermore, we show that solutions obtained by solving the convex problem can be improved by selecting the best of the models (for different regularization penalties) by using an criterion. In other words, an approximate solution to the right problem can be better than the exact solution to the wrong problem.
Keywords
Cite
@article{arxiv.1510.06319,
title = {A Risk Ratio Comparison of $l_0$ and $l_1$ Penalized Regression},
author = {Kory D. Johnson and Dongyu Lin and Lyle H. Ungar and Dean P. Foster and Robert A. Stine},
journal= {arXiv preprint arXiv:1510.06319},
year = {2015}
}