English
Related papers

Related papers: The Complete Lasso Tradeoff Diagram

200 papers

This paper provides two general classes of multiple decision functions where each member of the first class strongly controls the family-wise error rate (FWER), while each member of the second class strongly controls the false discovery…

Statistics Theory · Mathematics 2019-11-19 Edsel A. Pena , Joshua D. Habiger , Wensong Wu

Model-X knockoff has garnered significant attention among various feature selection methods due to its guarantees for controlling the false discovery rate (FDR). Since its introduction in parametric design, knockoff techniques have evolved…

Machine Learning · Computer Science 2024-11-11 Hongyu Shen , Yici Yan , Zhizhen Zhao

The problem of selecting a handful of truly relevant variables in supervised machine learning algorithms is a challenging problem in terms of untestable assumptions that must hold and unavailability of theoretical assurances that selection…

Methodology · Statistics 2023-11-10 Mehdi Rostami , Olli Saarela

There is a challenge in selecting high-dimensional mediators when the mediators have complex correlation structures and interactions. In this work, we frame the high-dimensional mediator selection problem into a series of hypothesis tests…

Methodology · Statistics 2025-09-16 Runqiu Wang , Ran Dai , Jieqiong Wang , Kah Meng Soh , Ziyang Xu , Mohamed Azzam , Hongying Dai , Cheng Zheng

The Lasso method is known to exhibit instability in the presence of highly correlated features, often leading to an arbitrary selection of predictors. This issue manifests itself in two primary error types: the erroneous omission of…

Methodology · Statistics 2025-08-07 Yanxin Liu , Yunqi Zhang

Inequalities are key tools to prove FDR control of a multiple test. The present paper studies upper and lower bounds for the FDR under various dependence structures of p-values, namely independence, reverse martingale dependence and…

Statistics Theory · Mathematics 2015-02-18 Philipp Heesen , Arnold Janssen

In traditional logistic regression models, the link function is often assumed to be linear and continuous in predictors. Here, we consider a threshold model that all continuous features are discretized into ordinal levels, which further…

Methodology · Statistics 2022-02-18 Yinan Lin , Wen Zhou , Zhi Geng , Gexin Xiao , Jianxin Yin

In classical statistics, the bias-variance trade-off describes how varying a model's complexity (e.g., number of fit parameters) affects its ability to make accurate predictions. According to this trade-off, optimal performance is achieved…

Machine Learning · Statistics 2022-08-05 Jason W. Rocks , Pankaj Mehta

The Lasso is one of the most ubiquitous methods for variable selection in high-dimensional linear regression and has been studied extensively under different regimes. In a particular asymptotic setup entailing $n/p\to \text{constant}$, an…

Statistics Theory · Mathematics 2026-02-10 Lina Hidmi , Asaf Weinstein

We develop a new class of distribution--free multiple testing rules for false discovery rate (FDR) control under general dependence. A key element in our proposal is a symmetrized data aggregation (SDA) approach to incorporating the…

Methodology · Statistics 2021-05-27 Lilun Du , Xu Guo , Wenguang Sun , Changliang Zou

We in this paper theoretically go over a rate-distortion based sparse dictionary learning problem. We show that the Degrees-of-Freedom (DoF) interested to be calculated $-$ satnding for the minimal set that guarantees our rate-distortion…

Information Theory · Computer Science 2022-06-15 Makan Zamanipour

Penalized regression methods, most notably the lasso, are a popular approach to analyzing high-dimensional data. An attractive property of the lasso is that it naturally performs variable selection. An important area of concern, however, is…

Methodology · Statistics 2026-05-13 Ryan Miller , Patrick Breheny

Public data repositories have enabled researchers to compare results across multiple genomic studies in order to replicate findings. A common approach is to first rank genes according to an hypothesis of interest within each study. Then,…

Applications · Statistics 2012-06-29 Loki Natarajan , Minya Pu , Karen Messer

Competition-based approach to controlling the false discovery rate (FDR) recently rose to prominence when, generalizing it to sequential hypothesis testing, Barber and Cand\`es used it as part of their knockoff-filter. Control of the FDR…

Methodology · Statistics 2023-02-24 Arya Ebadi , Dong Luo , Jack Freestone , William Stafford Noble , Uri Keich

High-dimensional sparse generalized linear models (GLMs) have emerged in the setting that the number of samples and the dimension of variables are large, and even the dimension of variables grows faster than the number of samples. False…

Statistics Theory · Mathematics 2021-05-04 Chang Cui , Jinzhu Jia , Yijun Xiao , Huiming Zhang

Federated Learning (FL) has recently received a lot of attention for large-scale privacy-preserving machine learning. However, high communication overheads due to frequent gradient transmissions decelerate FL. To mitigate the communication…

Machine Learning · Computer Science 2021-05-27 Milad Khademi Nori , Sangseok Yun , Il-Min Kim

False discovery rate (FDR) control is a popular approach for maintaining the integrity of statistical analyses, especially in high-dimensional data settings, where multiple comparisons increase the risk of false positives. FDR control has…

Signal Processing · Electrical Eng. & Systems 2026-03-03 Fabian Scheidt , Jasin Machkour , Michael Muma

Power and reproducibility are key to enabling refined scientific discoveries in contemporary big data applications with general high-dimensional nonlinear models. In this paper, we provide theoretical foundations on the power and robustness…

Statistics Theory · Mathematics 2017-09-04 Yingying Fan , Emre Demirkaya , Gaorong Li , Jinchi Lv

In modern scientific research, the objective is often to identify which variables are associated with an outcome among a large class of potential predictors. This goal can be achieved by selecting variables in a manner that controls the the…

Methodology · Statistics 2023-10-10 Yushu Shi , Michael Martens

False discovery rates (FDR) are an essential component of statistical inference, representing the propensity for an observed result to be mistaken. FDR estimates should accompany observed results to help the user contextualize the relevance…

Methodology · Statistics 2020-10-12 Megan Hollister Murray , Jeffrey D. Blume