English
Related papers

Related papers: Difference-in-differences Design with Outcomes Mis…

200 papers

Two key identifying assumptions used to justify difference-in-differences are parallel trends and no anticipation, yet both may fail in practice. I propose a class of assumptions on anticipation and derive closed-form, sharp bounds on the…

Econometrics · Economics 2026-03-03 Gianna Fenaroli

Treatment policy estimands are frequently favored by regulators, as they assess the effect of treatment assignment regardless of post-randomization events. Despite best efforts, missing data due to study discontinuation cannot be fully…

Methodology · Statistics 2026-05-13 Ajmal Oodally , Craig Wang , Zheng Li , Tim Morris , Tobias Mütze , Arunava Chakravartty

When one studies the effects of taxes, tariffs, or prices using panel data, the treatment is often continuously distributed in every period. We propose difference-in-differences (DID) estimators for such cases. We assume that between…

Training datasets for machine learning often have some form of missingness. For example, to learn a model for deciding whom to give a loan, the available training data includes individuals who were given a loan in the past, but not those…

Machine Learning · Computer Science 2020-12-22 Naman Goel , Alfonso Amayuelas , Amit Deshpande , Amit Sharma

While a difference-in-differences (DID) design was originally developed with one pre- and one post-treatment period, data from additional pre-treatment periods are often available. How can researchers improve the DID design with such…

Applications · Statistics 2022-02-14 Naoki Egami , Soichiro Yamauchi

When training predictive models on data with missing entries, the most widely used and versatile approach is a pipeline technique where we first impute missing entries and then compute predictions. In this paper, we view prediction with…

Machine Learning · Computer Science 2025-02-25 Dimitris Bertsimas , Arthur Delarue , Jean Pauphilet

The challenge of missing data remains a significant obstacle across various scientific domains, necessitating the development of advanced imputation techniques that can effectively address complex missingness patterns. This study introduces…

Machine Learning · Computer Science 2025-01-22 Harsh Joshi , Rajeshwari Mistri , Manasi Mali , Nachiket Kapure , Parul Kumari

Missing data is a pervasive problem in data analyses, resulting in datasets that contain censored realizations of a target distribution. Many approaches to inference on the target distribution using censored observed data, rely on missing…

Machine Learning · Statistics 2019-07-02 Rohit Bhattacharya , Razieh Nabi , Ilya Shpitser , James M. Robins

Missing data is a significant problem impacting all domains. State-of-the-art framework for minimizing missing data bias is multiple imputation, for which the choice of an imputation model remains nontrivial. We propose a multiple…

Machine Learning · Computer Science 2018-02-20 Lovedeep Gondara , Ke Wang

Understanding whether and how treatment effects vary across subgroups is crucial to inform clinical practice and recommendations. Accordingly, the assessment of heterogeneous treatment effects (HTE) based on pre-specified potential effect…

Methodology · Statistics 2023-12-04 Bryan S. Blette , Scott D. Halpern , Fan Li , Michael O. Harhay

This article proposes doubly robust estimators for the average treatment effect on the treated (ATT) in difference-in-differences (DID) research designs. In contrast to alternative DID estimators, the proposed estimators are consistent if…

Econometrics · Economics 2020-05-07 Pedro H. C. Sant'Anna , Jun B. Zhao

Missing values of varying patterns and rates in real-world tabular data pose a significant challenge in developing reliable data-driven models. The most commonly used statistical and machine learning methods for missing value imputation may…

Machine Learning · Computer Science 2025-03-26 Ibna Kowsar , Shourav B. Rabbani , Yina Hou , Manar D. Samad

In economic program evaluation, it is common to obtain panel data in which outcomes are indicators that an individual has reached an absorbing state. For example, they may indicate whether an individual has exited a period of unemployment,…

Econometrics · Economics 2026-05-26 Ben Deaner , Hyejin Ku

Assume that cause-effect relationships between variables can be described as a directed acyclic graph and the corresponding linear structural equation model.We consider the identification problem of total effects in the presence of latent…

Methodology · Statistics 2012-06-18 Zhihong Cai , Manabu Kuroki

Under what circumstances is it a threat to the parallel trends assumption required for Difference in Differences (DiD) studies if treatment decisions are based on past values of the outcome? We explore via simulation studies whether…

Methodology · Statistics 2022-08-02 Zach Shahn

Missing data are ubiquitous in many domains including healthcare. When these data entries are not missing completely at random, the (conditional) independence relations in the observed data may be different from those in the complete data…

Machine Learning · Computer Science 2020-07-14 Ruibo Tu , Kun Zhang , Paul Ackermann , Bo Christer Bertilson , Clark Glymour , Hedvig Kjellström , Cheng Zhang

Missing data is a common problem in clinical data collection, which causes difficulty in the statistical analysis of such data. To overcome problems caused by incomplete data, we propose a new imputation method called projective resampling…

Methodology · Statistics 2021-06-17 Zishu Zhan , Xiangjie Li , Jingxiao Zhang

Panel data analysis is an important topic in statistics and econometrics. Traditionally, in panel data analysis, all individuals are assumed to share the same unknown parameters, e.g. the same coefficients of covariates when the linear…

Statistics Theory · Mathematics 2017-06-09 Heng Lian , Xinghao Qiao , Wenyang Zhang

Missing time-series data is a prevalent practical problem. Imputation methods in time-series data often are applied to the full panel data with the purpose of training a model for a downstream out-of-sample task. For example, in finance,…

Machine Learning · Statistics 2023-04-13 Jose Blanchet , Fernando Hernandez , Viet Anh Nguyen , Markus Pelger , Xuhui Zhang

While a randomized control trial is considered the gold standard for estimating causal treatment effects, there are many research settings in which randomization is infeasible or unethical. In such cases, researchers rely on analytical…

Methodology · Statistics 2024-02-21 Julia C. Thome , Peter F. Rebeiro , Andrew J. Spieker , Bryan E. Shepherd