English
Related papers

Related papers: Who Are We Missing? A Principled Approach to Chara…

200 papers

Random forest regression (RF) is an extremely popular tool for the analysis of high-dimensional data. Nonetheless, its benefits may be lessened in sparse settings due to weak predictors, and a pre-estimation dimension reduction (targeting)…

Participants enrolled into randomized controlled trials (RCTs) often do not reflect real-world populations. Previous research in how best to translate RCT results to target populations has focused on weighting RCT data to look like the…

Augmenting randomized controlled trials (RCTs) with external real-world data (RWD) has the potential to improve the finite sample efficiency of treatment effect estimators. We describe using adaptive targeted maximum likelihood estimation…

Methodology · Statistics 2025-01-30 Sky Qiu , Jens Tarp , Andrew Mertens , Mark van der Laan

How can we effectively find the best structures in tree models? Tree models have been favored over complex black box models in domains where interpretability is crucial for making irreversible decisions. However, searching for a tree…

Machine Learning · Computer Science 2022-02-23 Jaemin Yoo , Lee Sael

Randomized controlled trials (RCTs) can be used to generate guarantees on treatment effects. However, RCTs often spend unnecessary resources exploring sub-optimal treatments, which can reduce the power of treatment guarantees. To address…

Computers and Society · Computer Science 2024-10-16 Santiago Cortes-Gomez , Naveen Raman , Aarti Singh , Bryan Wilder

Incorporating domain-specific constraints into machine learning models is essential for generating predictions that are both accurate and feasible in real-world applications. This paper introduces new methods for training Output-Constrained…

Machine Learning · Computer Science 2026-04-06 Hüseyin Tunç , Doğanay Özese , Ş. İlker Birbil , Donato Maragno , Marco Caserta , Mustafa Baydoğan

Randomized experiments have been used to assist decision-making in many areas. They help people select the optimal treatment for the test population with certain statistical guarantee. However, subjects can show significant heterogeneity in…

Artificial Intelligence · Computer Science 2017-05-25 Yan Zhao , Xiao Fang , David Simchi-Levi

Randomized controlled trials (RCTs) are considered the gold standard for estimating the average treatment effect (ATE) of interventions. One use of RCTs is to study the causes of global poverty -- a subject explicitly cited in the 2019…

Machine Learning · Computer Science 2023-05-26 Connor T. Jerzak , Fredrik Johansson , Adel Daoud

A common challenge for decision makers is selecting actions whose rewards are unknown and evolve over time based on prior policies. For instance, repeated use may reduce an action's effectiveness (habituation), while inactivity may restore…

Machine Learning · Computer Science 2025-11-06 Fengxu Li , Stephanie M. Carpenter , Matthew P. Buman , Yonatan Mintz

Randomized controlled trials (RCTs) frequently utilize covariate-adaptive randomization (CAR) (e.g., stratified block randomization) and commonly suffer from imperfect compliance. This paper studies the identification and inference for the…

Econometrics · Economics 2025-05-02 Federico A. Bugni , Mengsi Gao , Filip Obradovic , Amilcar Velez

Learning individualized treatment rules (ITRs) is an important topic in precision medicine. Current literature mainly focuses on deriving ITRs from a single source population. We consider the observational data setting when the source…

Machine Learning · Statistics 2023-07-04 Rui Chen , Jared D. Huling , Guanhua Chen , Menggang Yu

Tensor-on-tensor (TOT) regression is an important tool for the analysis of tensor data, aiming to predict a set of response tensors from a corresponding set of predictor tensors. However, standard TOT regression is sensitive to outliers,…

Methodology · Statistics 2026-03-30 Mehdi Hirari , Fabio Centofanti , Mia Hubert , Stefan Van Aelst

In any given machine learning problem, there may be many models that could explain the data almost equally well. However, most learning algorithms return only one of these models, leaving practitioners with no practical way to explore…

Machine Learning · Computer Science 2022-10-27 Rui Xin , Chudi Zhong , Zhi Chen , Takuya Takagi , Margo Seltzer , Cynthia Rudin

Randomized Controlled Trials (RCTs) may suffer from limited scope. In particular, samples may be unrepresentative: some RCTs over- or under- sample individuals with certain characteristics compared to the target population, for which one…

Methodology · Statistics 2024-03-15 Bénédicte Colnet , Julie Josse , Gaël Varoquaux , Erwan Scornet

Number sense is a core cognitive ability supporting various adaptive behaviors and is foundational for mathematical learning. Here, we study its emergence in unsupervised generative models through the lens of rate-distortion theory (RDT), a…

Neurons and Cognition · Quantitative Biology 2025-12-23 Leo D'Amato , Davide Nuzzi , Alberto Testolin , Ivilin Peev Stoianov , Marco Zorzi , Giovanni Pezzulo

Identifying and making statistical inferences on differential treatment effects (commonly known as subgroup analysis in clinical research) is central to precision health. Subgroup analysis allows practitioners to pinpoint populations for…

Machine Learning · Statistics 2026-02-05 Zhongming Xie , Joseph Giorgio , Jingshen Wang

This paper investigates an interesting weakly supervised regression setting called regression with interval targets (RIT). Although some of the previous methods on relevant regression settings can be adapted to RIT, they are not…

Machine Learning · Computer Science 2023-06-21 Xin Cheng , Yuzhou Cao , Ximing Li , Bo An , Lei Feng

Individualized treatment rules (ITR) can improve health outcomes by recognizing that patients may respond differently to treatment and assigning therapy with the most desirable predicted outcome for each individual. Flexible and efficient…

Methodology · Statistics 2017-09-25 Brent R. Logan , Rodney Sparapani , Robert E. McCulloch , Purushottam W. Laud

We present a method for incorporating missing data in non-parametric statistical learning without the need for imputation. We focus on a tree-based method, Bayesian Additive Regression Trees (BART), enhanced with "Missingness Incorporated…

Machine Learning · Statistics 2014-02-14 Adam Kapelner , Justin Bleich

Personalized medicine aims at identifying best treatments for a patient with given characteristics. It has been shown in the literature that these methods can lead to great improvements in medicine compared to traditional methods…

Machine Learning · Statistics 2018-11-26 Oleg Sysoev , Krzysztof Bartoszek , Eva-Charlotte Ekstrom , Katarina Ekholm Selling