English
Related papers

Related papers: Analysis of Two-Stage Rollout Designs with Cluster…

200 papers

Causal effect estimation in networked systems is central to data-driven decision making. In such settings, interventions on one unit can spill over to others, and in complex physical or social systems, the interaction pathways driving these…

Machine Learning · Statistics 2025-11-27 Sadegh Shirani , Mohsen Bayati

Machine learning systems increasingly depend on pipelines of multiple algorithms to provide high quality and well structured predictions. This paper argues interaction effects between clustering and prediction (e.g. classification,…

Machine Learning · Statistics 2019-01-01 Matt Barnes , Artur Dubrawski

In the first stage of a two-stage study, the researcher uses a statistical model to impute the unobserved exposures. In the second stage, imputed exposures serve as covariates in epidemiological models. Imputation error in the first stage…

Applications · Statistics 2021-07-19 Ron Sarafian , Itai Kloog , Jonathan D. Rosenblatt

This paper extends the design-based framework to settings with multi-way cluster dependence, and shows how multi-way clustering can be justified when clustered assignment and clustered sampling occurs on different dimensions, or when either…

Econometrics · Economics 2023-09-06 Luther Yap

Randomized saturation designs are a family of designs which assign a possibly different treatment proportion to each cluster of a population at random. As a result, they generalize the well-known (stratified) completely randomized designs…

Methodology · Statistics 2022-03-21 Chencheng Cai , Jean Pouget-Abadie , Edoardo M. Airoldi

Although there is now a large literature on policy evaluation and learning, much of the prior work assumes that the treatment assignment of one unit does not affect the outcome of another unit. Unfortunately, ignoring interference can lead…

Methodology · Statistics 2025-04-02 Yi Zhang , Kosuke Imai

We systematically investigate issues due to mis-specification that arise in estimating causal effects when (treatment) interference is informed by a network available pre-intervention, i.e., in situations where the outcome of a unit may…

Methodology · Statistics 2018-10-22 Vishesh Karwa , Edoardo M. Airoldi

The analysis of screening experiments is often done in two stages, starting with factor selection via an analysis under a main effects model. The success of this first stage is influenced by three components: (1) main effect estimators'…

Methodology · Statistics 2024-03-19 Jonathan W. Stallrich , Michael McKibben

Clustering is often used for discovering structure in data. Clustering systems differ in the objective function used to evaluate clustering quality and the control strategy used to search the space of clusterings. Ideally, the search…

Artificial Intelligence · Computer Science 2014-11-17 D. Fisher

Interference is ubiquitous when conducting causal experiments over networks. Except for certain network structures, causal inference on the network in the presence of interference is difficult due to the entanglement between the treatment…

Methodology · Statistics 2023-12-08 Chencheng Cai , Xu Zhang , Edoardo M. Airoldi

Increasingly, there is a marked interest in estimating causal effects under network interference due to the fact that interference manifests naturally in networked experiments. However, network information generally is available only up to…

Methodology · Statistics 2022-09-02 Wenrui Li , Daniel L. Sussman , Eric D. Kolaczyk

Causal inference analyses often use existing observational data, which in many cases has some clustering of individuals. In this paper we discuss propensity score weighting methods in a multilevel setting where within clusters individuals…

Applications · Statistics 2020-12-24 Youjin Lee , Trang Q. Nguyen , Elizabeth A. Stuart

We describe our framework, deployed at Facebook, that accounts for interference between experimental units through cluster-randomized experiments. We document this system, including the design and estimation procedures, and detail insights…

Social and Information Networks · Computer Science 2020-12-17 Brian Karrer , Liang Shi , Monica Bhole , Matt Goldman , Tyrone Palmer , Charlie Gelman , Mikael Konutgan , Feng Sun

Interference occurs when a unit's treatment (or exposure) affects another unit's outcome. In some settings, units may be grouped into clusters such that it is reasonable to assume that interference, if present, only occurs between…

Methodology · Statistics 2023-08-24 Chanhwa Lee , Donglin Zeng , Michael G. Hudgens

Split-plot designs find wide applicability in multifactor experiments with randomization restrictions. Practical considerations often warrant the use of unbalanced designs. This paper investigates randomization based causal inference in…

Methodology · Statistics 2019-06-21 Rahul Mukerjee , Tirthankar Dasgupta

Educational research often studies subjects that are in naturally clustered groups of classrooms or schools. When designing a randomized experiment to evaluate an intervention directed at teachers, but with effects on teachers and their…

Applications · Statistics 2009-08-17 Brenda Jenney , Sharon Lohr

Estimating causal effects on networks is challenging because treatments may affect both treated units and their neighbors, while network homophily induces dependence and confounding. These challenges are amplified when causal effects are…

Machine Learning · Statistics 2026-05-12 Yuanchen Wu , Yubai Yuan

We consider a potential outcomes model in which interference may be present between any two units but the extent of interference diminishes with spatial distance. The causal estimand is the global average treatment effect, which compares…

Methodology · Statistics 2022-09-16 Michael P. Leung

Cluster randomized trials (CRTs) randomly assign an intervention to groups of individuals (e.g., clinics or communities) and measure outcomes on individuals in those groups. While offering many advantages, this experimental design…

Cross-validation plays a fundamental role in Machine Learning, enabling robust evaluation of model performance and preventing overestimation on training and validation data. However, one of its drawbacks is the potential to create data…

Machine Learning · Computer Science 2025-08-28 Afonso Martini Spezia , Thomas Fontanari , Mariana Recamonde-Mendoza