English
Related papers

Related papers: A Cautionary Tale on Integrating Studies with Disp…

200 papers

This study introduces a data-driven, machine learning-based method to detect suitable control variables and instruments for assessing the causal effect of a treatment on an outcome in observational data. Our approach tests the joint…

Econometrics · Economics 2026-05-20 Nicolas Apfel , Julia Hatamyar , Martin Huber , Jannis Kueck

Consider the problem of estimating average treatment effects when a large number of covariates are used to adjust for possible confounding through outcome regression and propensity score models. The conventional approach of model building…

Statistics Theory · Mathematics 2018-01-31 Zhiqiang Tan

Measuring treatment effects in observational studies is challenging because of confounding bias. Confounding occurs when a variable affects both the treatment and the outcome. Traditional methods such as propensity score matching estimate…

Methodology · Statistics 2021-12-23 Bevan I. Smith , Charles Chimedza

Simulation studies are commonly used to evaluate the performance of newly developed meta-analysis methods. For methodology that is developed for an aggregated data meta-analysis, researchers often resort to simulation of the aggregated data…

Applications · Statistics 2022-01-19 Edwin R. van den Heuvel , Osama Almalik , Zhuozhao Zhan

How should researchers conduct causal inference when the outcome of interest is latent and measured imperfectly by multiple indicators? We develop a general nonparametric framework for identifying and estimating average treatment effects on…

Methodology · Statistics 2026-04-22 Jiawei Fu , Donald P. Green

A data science task can be deemed as making sense of the data or testing a hypothesis about it. The conclusions inferred from data can greatly guide us to make informative decisions. Big data has enabled us to carry out countless prediction…

Machine Learning · Computer Science 2022-01-12 Wenhao Zhang , Ramin Ramezani , Arash Naeim

When treating depression, clinicians are interested in determining the optimal treatment for a given patient, which is challenging given the amount of treatments available. To advance individualized treatment allocation, integrating data…

Individualized treatment decisions can improve health outcomes, but using data to make these decisions in a reliable, precise, and generalizable way is challenging with a single dataset. Leveraging multiple randomized controlled trials…

Making each modality in multi-modal data contribute is of vital importance to learning a versatile multi-modal model. Existing methods, however, are often dominated by one or few of modalities during model training, resulting in sub-optimal…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Yangyang Guo , Liqiang Nie , Harry Cheng , Zhiyong Cheng , Mohan Kankanhalli , Alberto Del Bimbo

In microbiome analysis, researchers often seek to identify taxonomic features associated with an outcome of interest. However, microbiome features are intercorrelated and linked by phylogenetic relationships, making it challenging to assess…

Methodology · Statistics 2023-09-18 Yushu Shi , Liangliang Zhang , Kim-Anh Do , Robert R. Jenq , Christine B. Peterson

A randomized trial and an analysis of observational data designed to emulate the trial sample observations separately, but have the same eligibility criteria, collect information on some shared baseline covariates, and compare the effects…

Methodology · Statistics 2022-03-29 Issa J. Dahabreh , Jon A. Steingrimsson , James M. Robins , Miguel A. Hernán

Integrating data from multiple heterogeneous sources has become increasingly popular to achieve a large sample size and diverse study population. This paper reviews development in causal inference methods that combines multiple datasets…

Methodology · Statistics 2021-10-05 Xu Shi , Ziyang Pan , Wang Miao

When there are multiple outcome series of interest, Synthetic Control analyses typically proceed by estimating separate weights for each outcome. In this paper, we instead propose estimating a common set of weights across outcomes, by…

Econometrics · Economics 2025-02-13 Liyang Sun , Eli Ben-Michael , Avi Feller

There is a growing need for flexible general frameworks that integrate individual-level data with external summary information for improved statistical inference. External information relevant for a risk prediction model may come in…

Methodology · Statistics 2023-04-11 Tian Gu , Jeremy M. G. Taylor , Bhramar Mukherjee

This paper focuses on developing Pareto-optimal estimation and policy learning to identify the most effective treatment that maximizes the total reward from both short-term and long-term effects, which might conflict with each other. For…

Machine Learning · Computer Science 2024-03-13 Yingrong Wang , Anpeng Wu , Haoxuan Li , Weiming Liu , Qiaowei Miao , Ruoxuan Xiong , Fei Wu , Kun Kuang

The difference-in-differences (DID) research design is a key identification strategy which allows researchers to estimate causal effects under the parallel trends assumption. While the parallel trends assumption is counterfactual and cannot…

Methodology · Statistics 2026-05-12 Jonas M. Mikhaeil , Christopher Harshaw

We study causal inference under case-control and case-population sampling. Specifically, we focus on the binary-outcome and binary-treatment case, where the parameters of interest are causal relative and attributable risks defined via the…

Econometrics · Economics 2023-10-24 Sung Jae Jun , Sokbae Lee

Combining observational and experimental data for causal inference can improve treatment effect estimation. However, many observational data sets cannot be released due to data privacy considerations, so one researcher may not have access…

Methodology · Statistics 2024-08-26 Charlotte Z. Mann , Adam C. Sales , Johann A. Gagnon-Bartsch

A Gaussian measurement error assumption, i.e., an assumption that the data are observed up to Gaussian noise, can bias any parameter estimation in the presence of outliers. A heavy tailed error assumption based on Student's t distribution…

Methodology · Statistics 2018-11-30 Hyungsuk Tak , Justin A. Ellis , Sujit K. Ghosh

We consider a problem of data integration. Consider determining which genes affect a disease. The genes, which we call predictor objects, can be measured in different experiments on the same individual. We address the question of finding…

Machine Learning · Statistics 2016-10-04 Xin Gao , Raymond J. Carroll