English
Related papers

Related papers: Conditional cross-fitting for unbiased machine-lea…

200 papers

We study the estimation of the ATE in randomized controlled trials under a dynamically evolving interference structure. This setting arises in applications such as ride-sharing, where drivers move over time, and social networks, where…

Statistics Theory · Mathematics 2025-11-11 Su Jia , Peter Frazier , Nathan Kallus , Christina Lee Yu

Existing weighting methods for treatment effect estimation are often built upon the idea of propensity scores or covariate balance. They usually impose strong assumptions on treatment assignment or outcome model to obtain unbiased…

Machine Learning · Computer Science 2023-05-09 Dongcheng Zhang , Kunpeng Zhang

The conditional average treatment effect (CATE) is the best measure of individual causal effects given baseline covariates. However, the CATE only captures the (conditional) average, and can overlook risks and tail events, which are…

Machine Learning · Statistics 2025-06-05 Nathan Kallus , Miruna Oprescu

While many areas of machine learning have benefited from the increasing availability of large and varied datasets, the benefit to causal inference has been limited given the strong assumptions needed to ensure identifiability of causal…

Machine Learning · Computer Science 2022-01-02 Wenshuo Guo , Serena Wang , Peng Ding , Yixin Wang , Michael I. Jordan

We present an optimized rerandomization design procedure for a non-sequential treatment-control experiment. Randomized experiments are the gold standard for finding causal effects in nature. But sometimes random assignments result in…

Methodology · Statistics 2021-01-26 Adam Kapelner , Abba M. Krieger , Michael Sklar , David Azriel

A core step of every algorithm for learning regression trees is the selection of the best splitting variable from the available covariates and the corresponding split point. Early tree algorithms (e.g., AID, CART) employed greedy search…

Methodology · Statistics 2019-06-26 Lisa Schlosser , Torsten Hothorn , Achim Zeileis

We propose generalized additive partial linear models for complex data which allow one to capture nonlinear patterns of some covariates, in the presence of linear components. The proposed method improves estimation efficiency and increases…

Statistics Theory · Mathematics 2014-05-26 Li Wang , Lan Xue , Annie Qu , Hua Liang

Controlled experiments are widely used in many applications to investigate the causal relationship between input factors and experimental outcomes. A completely randomized design is usually used to randomly assign treatment levels to…

Methodology · Statistics 2026-05-12 Yiou Li , Lulu Kang , Xiao Huang

A significant obstacle in the development of robust machine learning models is covariate shift, a form of distribution shift that occurs when the input distributions of the training and test sets differ while the conditional label…

Machine Learning · Statistics 2021-11-17 Nilesh Tripuraneni , Ben Adlam , Jeffrey Pennington

Although increasingly used for research, electronic health records (EHR) often lack gold-standard assessment of key data elements. Linking EHRs to other data sources with higher-quality measurements can improve statistical inference, but…

Methodology · Statistics 2025-03-05 Jenny Shen , Dane Isenberg , Kristin A. Linn , Rebecca A. Hubbard

There is growing interest in exploring causal effects in target populations via data combination. However, most approaches are tailored to specific settings and lack comprehensive comparative analyses. In this article, we focus on a typical…

Methodology · Statistics 2024-09-17 Peng Wu , Shanshan Luo , Zhi Geng

Classical randomized experiments, equipped with randomization-based inference, provide assumption-free inference for treatment effects. They have been the gold standard for drawing causal inference and provide excellent internal validity.…

Methodology · Statistics 2021-09-22 Zihao Yang , Tianyi Qu , Xinran Li

Numerous empirical studies employ regression discontinuity designs with multiple cutoffs and heterogeneous treatments. A common practice is to normalize all the cutoffs to zero and estimate one effect. This procedure identifies the average…

Econometrics · Economics 2021-01-06 Marinho Bertanha

Causal analyses for observational studies are often complicated by covariate imbalances among treatment groups, and matching methodologies alleviate this complication by finding subsets of treatment groups that exhibit covariate balance. It…

Methodology · Statistics 2021-04-26 Zach Branson

In randomized controlled trials (RCTs), treatment is often assigned by stratified randomization. I show that among all stratified randomization schemes which treat all units with probability one half, a certain matched-pair design achieves…

Econometrics · Economics 2022-06-17 Yuehao Bai

Observational studies provide invaluable opportunities to draw causal inference, but they may suffer from biases due to pretreatment difference between treated and control units. Matching is a popular approach to reduce observed covariate…

Methodology · Statistics 2025-09-17 Xinran Li

When an exposure of interest is confounded by unmeasured factors, an instrumental variable (IV) can be used to identify and estimate certain causal contrasts. Identification of the marginal average treatment effect (ATE) from IVs relies on…

Methodology · Statistics 2023-10-02 Alexander W. Levis , Matteo Bonvini , Zhenghao Zeng , Luke Keele , Edward H. Kennedy

Single-parameter summaries of variable effects in regression settings are desirable for ease of interpretation. However (partially) linear models for example, which would deliver these, may fit poorly to the data. On the other hand, an…

Statistics Theory · Mathematics 2025-07-28 Harvey Klyne , Rajen D. Shah

Understanding the dependencies among features of a dataset is at the core of most unsupervised learning tasks. However, a majority of generative modeling approaches are focused solely on the joint distribution $p(x)$ and utilize models…

Machine Learning · Computer Science 2020-08-07 Yang Li , Shoaib Akbar , Junier B. Oliva

High-dimensional compositional data arise naturally in many applications such as metagenomic data analysis. The observed data lie in a high-dimensional simplex, and conventional statistical methods often fail to produce sensible results due…

Methodology · Statistics 2016-01-19 Yuanpei Cao , Wei Lin , Hongzhe Li