English
Related papers

Related papers: On the Effect of Imputation on the 2SLS Variance

200 papers

Subsampling methods have been recently proposed to speed up least squares estimation in large scale settings. However, these algorithms are typically not robust to outliers or corruptions in the observed covariates. The concept of influence…

Machine Learning · Statistics 2014-06-20 Brian McWilliams , Gabriel Krummenacher , Mario Lucic , Joachim M. Buhmann

Electronic health records and other sources of observational data are increasingly used for drawing causal inferences. The estimation of a causal effect using these data not meant for research purposes is subject to confounding and…

Methodology · Statistics 2023-04-19 Janie Coulombe , Shu Yang

Despite its prevalence in statistical datasets, heteroscedasticity (non-constant sample variances) has been largely ignored in the high-dimensional statistics literature. Recently, studies have shown that the Lasso can accommodate…

Statistics Theory · Mathematics 2014-10-31 James Sharpnack , Mladen Kolar

In many empirical settings, directly observing a treatment variable may be infeasible although an error-prone surrogate measurement of the latter will often be available. Causal inference based solely on the surrogate measurement is…

Methodology · Statistics 2024-09-26 Ying Zhou , Eric Tchetgen Tchetgen

When using dyadic data (i.e., data indexed by pairs of units), researchers typically assume a linear model, estimate it using Ordinary Least Squares and conduct inference using ``dyadic-robust" variance estimators. The latter assumes that…

Econometrics · Economics 2024-11-20 Nathan Canen , Ko Sugiura

Data integration methods aim to extract low-dimensional embeddings from high-dimensional outcomes to remove unwanted variations, such as batch effects and unmeasured covariates, across heterogeneous datasets. However, multiple hypothesis…

Methodology · Statistics 2025-12-15 Jin-Hong Du , Kathryn Roeder , Larry Wasserman

Missing data is a common problem in medical research, and is commonly addressed using multiple imputation. Although traditional imputation methods allow for valid statistical inference when data are missing at random (MAR), their…

Classical causal and statistical inference methods typically assume the observed data consists of independent realizations. However, in many applications this assumption is inappropriate due to a network of dependences between units in the…

Machine Learning · Computer Science 2019-07-02 Rohit Bhattacharya , Daniel Malinsky , Ilya Shpitser

Noncompliance and missing data often occur in randomized trials, which complicate the inference of causal effects. When both noncompliance and missing data are present, previous papers proposed moment and maximum likelihood estimators for…

Methodology · Statistics 2014-09-04 Hua Chen , Peng Ding , Zhi Geng , Xiao-Hua Zhou

Missing exposure information is a very common feature of many observational studies. Here we study identifiability and efficient estimation of causal effects on vector outcomes, in such cases where treatment is unconfounded but partially…

Methodology · Statistics 2020-02-04 Edward H. Kennedy

Various methods have recently been proposed to estimate causal effects with confidence intervals that are uniformly valid over a set of data generating processes when high-dimensional nuisance models are estimated by post-model-selection or…

Methodology · Statistics 2025-10-07 Niloofar Moosavi , Tetiana Gorbach , Xavier de Luna

To perform regression analysis in high dimensions, lasso or ridge estimation are a common choice. However, it has been shown that these methods are not robust to outliers. Therefore, alternatives as penalized M-estimation or the sparse…

Statistics Theory · Mathematics 2025-02-03 Viktoria Öllerer , Christophe Croux , Andreas Alfons

We propose a new 2-stage procedure that relies on the elastic net penalty to estimate a network based on partial correlations when data are heavy-tailed. The new estimator allows to consider the lasso penalty as a special case. Using Monte…

Methodology · Statistics 2021-08-25 Davide Bernardini , Sandra Paterlini , Emanuele Taufer

We study the asymptotic properties of the GLS estimator in multivariate regression with heteroskedastic and autocorrelated errors. We derive Wald statistics for linear restrictions and assess their performance. The statistics remains robust…

Econometrics · Economics 2025-03-19 Koichiro Moriya , Akihiko Noda

We investigate the problem of statistical inference for logistic regression with high-dimensional covariates in settings where dependence among individuals is induced by an underlying Markov random field. Going beyond the pairwise…

Statistics Theory · Mathematics 2026-03-23 Josh Miles , Sohom Bhattacharya

We develop a method to perform model averaging in two-stage linear regression systems subject to endogeneity. Our method extends an existing Gibbs sampler for instrumental variables to incorporate a component of model uncertainty. Direct…

Methodology · Statistics 2012-03-20 Anna Karl , Alex Lenkoski

The two-stage least-squares (2SLS) estimator is known to be biased when its first-stage fit is poor. I show that better first-stage prediction can alleviate this bias. In a two-stage linear regression model with Normal noise, I consider…

Statistics Theory · Mathematics 2017-11-01 Jann Spiess

To estimate the causal effects of beliefs on actions, researchers often run information provision experiments. We consider the causal interpretation of two-stage least squares (TSLS) estimators in these experiments. We characterize common…

Econometrics · Economics 2024-06-24 Vod Vilfort , Whitney Zhang

This work extends causal inference with stochastic confounders. We propose a new approach to variational estimation for causal inference based on a representer theorem with a random input space. We estimate causal effects involving latent…

Machine Learning · Statistics 2021-01-26 Thanh Vinh Vo , Pengfei Wei , Wicher Bergsma , Tze-Yun Leong

Estimating the causal effects of an intervention from high-dimensional observational data is difficult due to the presence of confounding. The task is often complicated by the fact that we may have a systematic missingness in our data at…

Machine Learning · Statistics 2020-03-02 Sonali Parbhoo , Mario Wieser , Aleksander Wieczorek , Volker Roth