English

Signed Support Recovery for Single Index Models in High-Dimensions

Statistics Theory 2016-06-24 v2 Machine Learning Statistics Theory

Abstract

In this paper we study the support recovery problem for single index models Y=f(Xβ,ε)Y=f(\boldsymbol{X}^{\intercal} \boldsymbol{\beta},\varepsilon), where ff is an unknown link function, XNp(0,Ip)\boldsymbol{X}\sim N_p(0,\mathbb{I}_{p}) and β\boldsymbol{\beta} is an ss-sparse unit vector such that βi{±1s,0}\boldsymbol{\beta}_{i}\in \{\pm\frac{1}{\sqrt{s}},0\}. In particular, we look into the performance of two computationally inexpensive algorithms: (a) the diagonal thresholding sliced inverse regression (DT-SIR) introduced by Lin et al. (2015); and (b) a semi-definite programming (SDP) approach inspired by Amini & Wainwright (2008). When s=O(p1δ)s=O(p^{1-\delta}) for some δ>0\delta>0, we demonstrate that both procedures can succeed in recovering the support of β\boldsymbol{\beta} as long as the rescaled sample size κ=nslog(ps)\kappa=\frac{n}{s\log(p-s)} is larger than a certain critical threshold. On the other hand, when κ\kappa is smaller than a critical value, any algorithm fails to recover the support with probability at least 12\frac{1}{2} asymptotically. In other words, we demonstrate that both DT-SIR and the SDP approach are optimal (up to a scalar) for recovering the support of β\boldsymbol{\beta} in terms of sample size. We provide extensive simulations, as well as a real dataset application to help verify our theoretical observations.

Cite

@article{arxiv.1511.02270,
  title  = {Signed Support Recovery for Single Index Models in High-Dimensions},
  author = {Matey Neykov and Qian Lin and Jun S. Liu},
  journal= {arXiv preprint arXiv:1511.02270},
  year   = {2016}
}

Comments

38 pages, 7 figures; 1 table; data set analysis added; typos corrected

R2 v1 2026-06-22T11:39:28.013Z