English
Related papers

Related papers: Debiased Lasso After Sample Splitting for Estimati…

200 papers

We propose a new approach to safe variable preselection in high-dimensional penalized regression, such as the lasso. Preselection - to start with a manageable set of covariates - has often been implemented without clear appreciation of its…

Semi-supervised learning by self-training heavily relies on pseudo-label selection (PLS). The selection often depends on the initial model fit on labeled data. Early overfitting might thus be propagated to the final model by selecting…

Machine Learning · Statistics 2023-06-27 Julian Rodemann , Jann Goschenhofer , Emilio Dorigatti , Thomas Nagler , Thomas Augustin

We propose an L1-penalized algorithm for fitting high-dimensional generalized linear mixed models. Generalized linear mixed models (GLMMs) can be viewed as an extension of generalized linear models for clustered observations. This…

Computation · Statistics 2014-06-03 Jürg Schelldorfer , Lukas Meier , Peter Bühlmann

Random sampling is a fundamental tool in modern machine learning and numerical linear algebra for reducing the computational cost of large-scale matrix problems. Existing analyses, however, rely primarily on subspace embedding guarantees,…

Numerical Analysis · Mathematics 2026-05-26 Chengmei Niu , Sachin Garg , Michał Dereziński , Zhenyu Liao

This paper proposes a debiased estimator for causal effects in high-dimensional generalized linear models with binary outcomes and general link functions. The estimator augments a regularized regression plug-in with weights computed from a…

Econometrics · Economics 2025-10-21 Jing Kong

We introduce a two-stage probabilistic framework for statistical downscaling using unpaired data. Statistical downscaling seeks a probabilistic map to transform low-resolution data from a biased coarse-grained numerical scheme to…

Machine Learning · Computer Science 2023-11-01 Zhong Yi Wan , Ricardo Baptista , Yi-fan Chen , John Anderson , Anudhyan Boral , Fei Sha , Leonardo Zepeda-Núñez

Simultaneous inference for high-dimensional non-Gaussian time series is always considered to be a challenging problem. Such tasks require not only robust estimation of the coefficients in the random process, but also deriving limiting…

Methodology · Statistics 2021-11-03 Linbo Liu , Danna Zhang

Classically, confidence intervals are required to have consistent coverage across all values of the parameter. However, this will inevitably break down if the underlying estimation procedure is biased. For this reason, many efforts have…

Methodology · Statistics 2025-08-06 Logan Harris , Patrick Breheny

Effect modification occurs when the effect of the treatment on an outcome varies according to the level of other covariates and often has important implications in decision making. When there are tens or hundreds of covariates, it becomes…

Methodology · Statistics 2021-11-23 Qingyuan Zhao , Dylan S. Small , Ashkan Ertefaie

Fitting statistical models is computationally challenging when the sample size or the dimension of the dataset is huge. An attractive approach for down-scaling the problem size is to first partition the dataset into subsets and then fit…

Methodology · Statistics 2016-02-15 Xiangyu Wang , David Dunson , Chenlei Leng

Post-Double-Lasso is becoming the most popular method for estimating linear regression models with many covariates when the purpose is to obtain an accurate estimate of a parameter of interest, such as an average treatment effect. However,…

Econometrics · Economics 2025-11-27 Sullivan Hué , Sébastien Laurent , Ulrich Aiounou , Emmanuel Flachaire

Nowadays an increasing amount of data is available and we have to deal with models in high dimension (number of covariates much larger than the sample size). Under sparsity assumption it is reasonable to hope that we can make a good…

Statistics Theory · Mathematics 2014-01-23 Mélanie Blazère , Jean-Michel Loubes , Fabrice Gamboa

These lecture notes consist of three chapters. In the first chapter we present oracle inequalities for the prediction error of the Lasso and square-root Lasso and briefly describe the scaled Lasso. In the second chapter we establish…

Statistics Theory · Mathematics 2014-10-01 Sara van de Geer

An informative sampling design leads to the selection of units whose inclusion probabilities are correlated with the response variable of interest. Model inference performed on the resulting observed sample will be biased for the population…

Methodology · Statistics 2018-06-29 Matthew R. Williams , Terrance D. Savitsky

We provide a principled way for investigators to analyze randomized experiments when the number of covariates is large. Investigators often use linear multivariate regression to analyze randomized experiments instead of simply reporting the…

Statistics Theory · Mathematics 2022-06-08 Adam Bloniarz , Hanzhong Liu , Cun-Hui Zhang , Jasjeet Sekhon , Bin Yu

We study the finite sample behavior of Lasso-based inference methods such as post double Lasso and debiased Lasso. We show that these methods can exhibit substantial omitted variable biases (OVBs) due to Lasso not selecting relevant…

Statistics Theory · Mathematics 2021-09-15 Kaspar Wuthrich , Ying Zhu

We present large sample results for partitioning-based least squares nonparametric regression, a popular method for approximating conditional expectation functions in statistics, econometrics, and machine learning. First, we obtain a…

Statistics Theory · Mathematics 2020-07-20 Matias D. Cattaneo , Max H. Farrell , Yingjie Feng

Selection of covariates is crucial in the estimation of average treatment effects given observational data with high or even ultra-high dimensional pretreatment variables. Existing methods for this problem typically assume sparse linear…

Methodology · Statistics 2023-03-20 Juan Chen , Yingchun Zhou

rdering of regression or classification coefficients occurs in many real-world applications. Fused Lasso exploits this ordering by explicitly regularizing the differences between neighboring coefficients through an $\ell_1$ norm…

Computation · Statistics 2010-06-29 Gui-Bo Ye , Xiaohui Xie

We study the problem of signal estimation from non-linear observations when the signal belongs to a low-dimensional set buried in a high-dimensional space. A rough heuristic often used in practice postulates that non-linear observations may…

Information Theory · Computer Science 2015-11-17 Yaniv Plan , Roman Vershynin