English
Related papers

Related papers: Demystify Doubly-Robust Estimation: The Role of Ov…

200 papers

Covariate imbalance between treatment groups makes it difficult to compare cumulative incidence curves in competing risk analyses. In this paper we discuss different methods to estimate adjusted cumulative incidence curves including inverse…

Difference-in-differences (DID) is a widely used approach for drawing causal inference from observational panel data. Two common estimation strategies for DID are outcome regression and propensity score weighting. In this paper, motivated…

Applications · Statistics 2021-01-05 Fan Li , Fan Li

Overlap, also known as positivity, is a key condition for causal treatment effect estimation. Many popular estimators suffer from high variance and become brittle when features differ strongly across treatment groups. This is especially…

Machine Learning · Statistics 2026-04-02 Oscar Clivio , Alexander D'Amour , Alexander Franks , David Bruns-Smith , Chris Holmes , Avi Feller

The consistency of doubly robust estimators relies on consistent estimation of at least one of two nuisance regression parameters. In moderate to large dimensions, the use of flexible data-adaptive regression estimators may aid in achieving…

Machine Learning · Statistics 2019-01-30 Iván Díaz

The weighted average treatment effect (WATE) defines a versatile class of causal estimands for populations characterized by propensity score weights, including the average treatment effect (ATE), treatment effect on the treated (ATT), on…

Methodology · Statistics 2025-09-23 Yiming Wang , Yi Liu , Shu Yang

Micro-randomized trials are commonly conducted for optimizing mobile health interventions such as push notifications for behavior change. In analyzing such trials, causal excursion effects are often of primary interest, and their estimation…

Methodology · Statistics 2024-08-19 Yihan Bao , Lauren Bell , Elizabeth Williamson , Claire Garnett , Tianchen Qian

This paper investigates the problem of making inference about a parametric model for the regression of an outcome variable $Y$ on covariates $(V,L)$ when data are fused from two separate sources, one which contains information only on $(V,…

Methodology · Statistics 2020-12-15 Katherine Evans , BaoLuo Sun , James Robins , Eric J. Tchetgen Tchetgen

Longitudinal data often involve heterogeneity, sparse signals, and contamination from response outliers or high-leverage observations especially in biomedical science. Existing methods usually address only part of this problem, either…

Methodology · Statistics 2026-02-26 Yuyao Wang , Yu Lu , Tianni Zhang , Mengfei Ran

Estimating dynamic treatment effects is a crucial endeavor in causal inference, particularly when confronted with high-dimensional confounders. Doubly robust (DR) approaches have emerged as promising tools for estimating treatment effects…

Methodology · Statistics 2023-05-17 Jelena Bradic , Weijie Ji , Yuqian Zhang

In this paper, I try to tame "Basu's elephants" (data with extreme selection on observables). I propose new practical large-sample and finite-sample methods for estimating and inferring heterogeneous causal effects (under unconfoundedness)…

Econometrics · Economics 2023-01-20 Ganesh Karapakula

The Doubly Robust (DR) estimation of ATE can be carried out in 2 steps, where in the first step, the treatment and outcome are modeled, and in the second step the predictions are inserted into the DR estimator. The model misspecification in…

Methodology · Statistics 2021-08-03 Mehdi Rostami , Olli Saarela , Michael Escobar

In prevalent cohort studies with follow-up, the time-to-event outcome is subject to left truncation leading to selection bias. For estimation of the distribution of time-to-event, conventional methods adjusting for left truncation tend to…

Methodology · Statistics 2025-12-29 Yuyao Wang , Andrew Ying , Ronghui Xu

Importance-weighting is a popular and well-researched technique for dealing with sample selection bias and covariate shift. It has desirable characteristics such as unbiasedness, consistency and low computational complexity. However,…

Machine Learning · Statistics 2019-03-12 Wouter M. Kouw , Marco Loog

In this expository note we describe a surprising phenomenon in overparameterized linear regression, where the dimension exceeds the number of samples: there is a regime where the test risk of the estimator found by gradient descent…

Machine Learning · Statistics 2019-12-17 Preetum Nakkiran

We study off-policy evaluation in the setting of contextual bandits, where we aim to evaluate a new policy using historical data that consists of contexts, actions and received rewards. This historical data typically does not faithfully…

Machine Learning · Computer Science 2026-03-11 Rong J. B. Zhu

This article develops a covariate balancing approach for the estimation of treatment effects on the treated (ATT) in a difference-in-differences (DID) research design when panel data are available. We show that the proposed covariate…

Econometrics · Economics 2025-08-05 Junjie Li , Yukitoshi Matsushita

In empirical studies with time-to-event outcomes, investigators often leverage observational data to conduct causal inference on the effect of exposure when randomized controlled trial data is unavailable. Model misspecification and lack of…

Methodology · Statistics 2023-05-05 Shenbo Xu , Bang Zheng , Bowen Su , Stan Finkelstein , Roy Welsch , Kenney Ng , Ioanna Tzoulaki , Zach Shahn

Randomized Controlled Trials (RCTs) may suffer from limited scope. In particular, samples may be unrepresentative: some RCTs over- or under- sample individuals with certain characteristics compared to the target population, for which one…

Methodology · Statistics 2024-03-15 Bénédicte Colnet , Julie Josse , Gaël Varoquaux , Erwan Scornet

There are now many options for doubly robust estimation; however, there is a concerning trend in the applied literature to believe that the combination of a propensity score and an adjusted outcome model automatically results in a doubly…

We consider the problem of estimating the finite population mean $\bar{Y}$ of an outcome variable $Y$ using data from a nonprobability sample and auxiliary information from a probability sample. Existing double robust (DR) estimators of…

Methodology · Statistics 2025-10-30 Shaun Seaman