English
Related papers

Related papers: Causal Inference from Possibly Unbalanced Split-Pl…

200 papers

This paper focuses on the design of spatial experiments to optimize the amount of information derived from the experimental data and enhance the accuracy of the resulting causal effect estimator. We propose a surrogate function for the mean…

Machine Learning · Computer Science 2025-08-29 Jin Zhu , Jingyi Li , Hongyi Zhou , Yinan Lin , Zhenhua Lin , Chengchun Shi

We show how clustering standard errors in one or more dimensions can be justified in M-estimation when there is sampling or assignment uncertainty. Since existing procedures for variance estimation are either conservative or invalid, we…

Econometrics · Economics 2024-11-21 Ruonan Xu , Luther Yap

Transductive graph-based semi-supervised learning methods usually build an undirected graph utilizing both labeled and unlabeled samples as vertices. Those methods propagate label information of labeled samples to neighbors through their…

Machine Learning · Computer Science 2013-12-25 Fengqi Li , Chuang Yu , Nanhai Yang , Feng Xia , Guangming Li , Fatemeh Kaveh-Yazdy

Importance sampling is often used in machine learning when training and testing data come from different distributions. In this paper we propose a new variant of importance sampling that can reduce the variance of importance sampling-based…

Machine Learning · Computer Science 2016-11-11 Philip S. Thomas , Emma Brunskill

Predicting the effect of interventions with many possible variations, e.g., therapeutic content that affects mental health outcomes or an earnings call transcript that drives movement in share price, is useful across several domains.…

Machine Learning · Computer Science 2026-05-27 Nikita Dhawan , Arnav Paruthi , Andrew Kim , Lovedeep Gondara , Jekaterina Novikova , Chris J. Maddison

In contrast to problems of interference in (exogenous) treatments, models of interference in unit-specific (endogenous) outcomes do not usually produce a reduced-form representation where outcomes depend on other units' treatment status…

Econometrics · Economics 2025-06-17 Konrad Menzel

We consider the problem of inference in shift-share research designs. The choice between existing approaches that allow for unrestricted spatial correlation involves tradeoffs, varying in terms of their validity when there are relatively…

Econometrics · Economics 2022-06-03 Luis Alvarez , Bruno Ferman , Raoni Oliveira

Paired cluster-randomized experiments (pCRTs) are common across many disciplines because there is often natural clustering of individuals, and paired randomization can help balance baseline covariates to improve experimental precision.…

Methodology · Statistics 2024-07-03 Charlotte Z. Mann , Adam C. Sales , Johann A. Gagnon-Bartsch

Causal inference from observational data requires assumptions. These assumptions range from measuring confounders to identifying instruments. Traditionally, causal inference assumptions have focused on estimation of effects for a single…

Machine Learning · Statistics 2019-03-04 Rajesh Ranganath , Adler Perotte

Randomized experiments are the gold standard for estimating the causal effects of an intervention. In the simplest setting, each experimental unit is randomly assigned to receive treatment or control, and then the outcomes in each treatment…

Methodology · Statistics 2020-06-05 Guillaume Basse , Yi Ding , Panos Toulis

A design-based individual prediction approach is developed based on the expected cross-validation results, given the sampling design and the sample-splitting design for cross-validation. Whether the predictor is selected from an ensemble of…

Machine Learning · Statistics 2023-01-24 Li-Chun Zhang , Danhyang Lee

While many areas of machine learning have benefited from the increasing availability of large and varied datasets, the benefit to causal inference has been limited given the strong assumptions needed to ensure identifiability of causal…

Machine Learning · Computer Science 2022-01-02 Wenshuo Guo , Serena Wang , Peng Ding , Yixin Wang , Michael I. Jordan

We study the problem of estimating causal effects under hidden confounding in the following unpaired data setting: we observe some covariates $X$ and an outcome $Y$ under different experimental conditions (environments) but do not observe…

Machine Learning · Statistics 2026-01-22 Felix Schur , Niklas Pfister , Peng Ding , Sach Mukherjee , Jonas Peters

While attractive from a theoretical perspective, finely stratified experiments such as paired designs suffer from certain analytical limitations not present in block-randomized experiments with multiple treated and control individuals in…

Methodology · Statistics 2017-06-21 Colin B. Fogarty

Conservative inference is a major concern in simulation-based inference. It has been shown that commonly used algorithms can produce overconfident posterior approximations. Balancing has empirically proven to be an effective way to mitigate…

Machine Learning · Statistics 2023-04-24 Arnaud Delaunoy , Benjamin Kurt Miller , Patrick Forré , Christoph Weniger , Gilles Louppe

Experiments using multi-step protocols often involve several restrictions on the randomization. For a specific application to in vitro testing on microplates, a design was required with both a split-plot and a strip-plot structure. On top…

We consider the problem of how to assign treatment in a randomized experiment, in which the correlation among the outcomes is informed by a network available pre-intervention. Working within the potential outcome causal framework, we…

Methodology · Statistics 2017-05-19 Guillaume W. Basse , Edoardo M. Airoldi

Recently, interest has grown in the use of proxy variables of unobserved confounding for inferring the causal effect in the presence of unmeasured confounders from observational data. One difficulty inhibiting the practical use is finding…

Machine Learning · Computer Science 2024-05-28 Feng Xie , Zhengming Chen , Shanshan Luo , Wang Miao , Ruichu Cai , Zhi Geng

Randomized experiments are the "gold standard" for estimating causal effects, yet often in practice, chance imbalances exist in covariate distributions between treatment groups. If covariate data are available before units are exposed to…

Statistics Theory · Mathematics 2012-07-25 Kari Lock Morgan , Donald B. Rubin

Pooling multiple neuroimaging datasets across institutions often enables improvements in statistical power when evaluating associations (e.g., between risk factors and disease outcomes) that may otherwise be too weak to detect. When there…

Machine Learning · Computer Science 2022-03-30 Vishnu Suresh Lokhande , Rudrasis Chakraborty , Sathya N. Ravi , Vikas Singh