English

Controlling the false discovery rate in high-dimensional linear models using model-X knockoffs and $p$-values

Methodology 2026-03-17 v2 Statistics Theory Statistics Theory

Abstract

We propose a novel multiple testing methodology for controlling the false discovery rate (FDR) in high-dimensional linear models that integrates model-X knockoff techniques with debiased penalized regression estimators. At the foundation of our methodology, we construct and study two sets of naturally paired high-dimensional test statistics and the associated pp-values for evaluating the same null hypotheses. The first set is shown to be asymptotically mutually independent, justifying the use of the Benjamini-Hochberg procedure. We further exploit the pairing structure through a two-step procedure aimed at improving power. Our theoretical results establish the key properties of the framework with respect to asymptotic FDR control and formally characterize the associated power gains of the two-step procedure. Importantly, our framework accommodates general dependence in the design matrix. Extensive simulations demonstrate that our methods outperform existing approaches -- particularly those relying on empirical FDP estimates -- in both power and FDR control accuracy, with notable gains in settings involving weaker signals, small sample sizes, or low target FDR levels.

Keywords

Cite

@article{arxiv.2505.16124,
  title  = {Controlling the false discovery rate in high-dimensional linear models using model-X knockoffs and $p$-values},
  author = {Jinyuan Chang and Chenlong Li and Cheng Yong Tang and Zhengtian Zhu},
  journal= {arXiv preprint arXiv:2505.16124},
  year   = {2026}
}
R2 v1 2026-07-01T02:30:08.137Z