English
Related papers

Related papers: On Varieties of Doubly Robust Estimators Under Mis…

200 papers

Valid estimation of treatment effects from observational data requires proper control of confounding. If the number of covariates is large relative to the number of observations, then controlling for all available covariates is infeasible.…

Methodology · Statistics 2018-01-11 Joseph Antonelli , Matthew Cefalu , Nathan Palmer , Denis Agniel

Due to concerns about parametric model misspecification, there is interest in using machine learning to adjust for confounding when evaluating the causal effect of an exposure on an outcome. Unfortunately, exposure effect estimators that…

Methodology · Statistics 2025-01-08 Oliver Dukes , Stijn Vansteelandt , David Whitney

Missing exposure information is a very common feature of many observational studies. Here we study identifiability and efficient estimation of causal effects on vector outcomes, in such cases where treatment is unconfounded but partially…

Methodology · Statistics 2020-02-04 Edward H. Kennedy

Predictors are learned using past training data which may contain features that are unavailable at the time of prediction. We develop an approach that is robust against outlying missing features, based on the optimality properties of an…

Signal Processing · Electrical Eng. & Systems 2020-07-15 Xiuming Liu , Dave Zachariah , Petre Stoica

We present a second-order estimator of the mean of a variable subject to missingness, under the missing at random assumption. The estimator improves upon existing methods by using an approximate second-order expansion of the parameter…

Statistics Theory · Mathematics 2015-11-30 Iván Díaz , Marco Carone , Mark J. van der Laan

While model selection is a well-studied topic in parametric and nonparametric regression or density estimation, selection of possibly high-dimensional nuisance parameters in semiparametric problems is far less developed. In this paper, we…

Methodology · Statistics 2023-09-06 Yifan Cui , Eric Tchetgen Tchetgen

Missing data in supervised learning is well-studied, but the specific issue of missing labels during model evaluation has been overlooked. Ignoring samples with missing values, a common solution, can introduce bias, especially when data is…

Machine Learning · Computer Science 2025-04-28 Danial Dervovic , Michael Cashmore

Integrating probability and nonprobability survey samples is an important problem in modern survey sampling. Nonprobability samples often contain rich outcome information but may lack population representativeness, whereas probability…

Statistics Theory · Mathematics 2026-05-28 Yufang Dai , Shihua Luo , Wendy Lou , Zilin Wang , Xuewen Lu

We study the problem of off-policy value evaluation in reinforcement learning (RL), where one aims to estimate the value of a new policy based on data collected by a different policy. This problem is often a critical step when applying RL…

Machine Learning · Computer Science 2016-05-27 Nan Jiang , Lihong Li

Missing covariates are not uncommon in capture-recapture studies. When covariate information is missing at random in capture-recapture data, an empirical full likelihood method has been demonstrated to outperform…

Methodology · Statistics 2025-07-15 Yang Liu , Yukun Liu , Pengfei Li , Riquan Zhang

Let $X$ be a random variable with unknown mean and finite variance. We present a new estimator of the mean of $X$ that is robust with respect to the possible presence of outliers in the sample, provides tight sub-Gaussian deviation…

Statistics Theory · Mathematics 2022-01-03 Stanislav Minsker , Mohamed Ndaoud

Robust estimators of large covariance matrices are considered, comprising regularized (linear shrinkage) modifications of Maronna's classical M-estimators. These estimators provide robustness to outliers, while simultaneously being…

Statistics Theory · Mathematics 2018-07-04 Nicolas Auguin , David Morales-Jimenez , Matthew R. McKay , Romain Couillet

Doubly robust (DR) estimators guard against model misspecification but remain sensitive to weak covariate overlap. We show that trimming propensity scores reduces variance but eliminates double robustness. We introduce DR estimators that…

Econometrics · Economics 2026-04-17 Yukun Ma , Pedro H. C. Sant'Anna , Yuya Sasaki , Takuya Ura

The marginal structure quantile model (MSQM) provides a unique lens to understand the causal effect of a time-varying treatment on the full distribution of potential outcomes. Under the semiparametric framework, we derive the efficiency…

Methodology · Statistics 2024-02-13 Chao Cheng , Liangyuan Hu , Fan Li

In this paper we study covariance estimation with missing data. We consider missing data mechanisms that can be independent of the data, or have a time varying dependency. Additionally, observed variables may have arbitrary (non uniform)…

Statistics Theory · Mathematics 2021-06-17 Eduardo Pavez , Antonio Ortega

We study counterfactual regression, which aims to map input features to outcomes under hypothetical scenarios that differ from those observed in the data. This is particularly useful for decision-making when adapting to sudden shifts in…

Methodology · Statistics 2025-04-08 Kwangho Kim

The research described herewith is to re-visit the classical doubly robust estimation of average treatment effect by conducting a systematic study on the comparisons, in the sense of asymptotic efficiency, among all possible combinations of…

Statistics Theory · Mathematics 2020-06-01 Keli Guo , Chuyun Ye , Jun Fan , Lixing Zhu

The Difference-in-Differences (DiD) method is a fundamental tool for causal inference, yet its application is often complicated by missing data. Although recent work has developed robust DiD estimators for complex settings like staggered…

Methodology · Statistics 2026-01-27 Lorenzo Testa , Edward H. Kennedy , Matthew Reimherr

We study counterfactual classification as a new tool for decision-making under hypothetical (contrary to fact) scenarios. We propose a doubly-robust nonparametric estimator for a general counterfactual classifier, where we can incorporate…

Machine Learning · Computer Science 2023-01-31 Kwangho Kim , Edward H. Kennedy , José R. Zubizarreta

In a missing-data setting, we have a sample in which a vector of explanatory variables x_i is observed for every subject i, while scalar outcomes y_i are missing by happenstance on some individuals. In this work we propose robust estimates…

Statistics Theory · Mathematics 2010-09-20 Mariela Sued , Victor J. Yohai