English
Related papers

Related papers: A synthetic data integration framework to leverage…

200 papers

Multiple imputation has become one of the standard methods in drawing inferences in many incomplete data applications. Applications of multiple imputation in relatively more complex settings, such as high-dimensional clustered data, require…

Methodology · Statistics 2025-04-08 Qiushuang Li , Recai Yucel

Background: Pairwise and network meta-analyses using fixed effect and random effects models are commonly applied to synthesise evidence from randomised controlled trials. The models differ in their assumptions and the interpretation of the…

Methodology · Statistics 2017-08-04 Shijie Ren , Jeremy E. Oakley , John W. Stevens

We develop a generally applicable full-information inference method for heterogeneous agent models, combining aggregate time series data and repeated cross sections of micro data. To handle unobserved aggregate state variables that affect…

Econometrics · Economics 2024-11-19 Laura Liu , Mikkel Plagborg-Møller

Suppose we have individual data from an internal study and various summary statistics from relevant external studies. External summary statistics have the potential to improve statistical inference for the internal population; however, it…

Methodology · Statistics 2026-02-06 Wenjie Hu , Ruoyu Wang , Wei Li , Wang Miao

In many modern applications, a carefully designed primary study provides individual-level data for interpretable modeling, while summary-level external information is available through black-box, efficient, and nonparametric…

Methodology · Statistics 2026-04-07 Chi-Shian Dai , Jun Shao

Survival regression is widely used to model time-to-events data, to explore how covariates may influence the occurrence of events. Modern datasets often encompass a vast number of covariates across many subjects, with only a subset of the…

Methodology · Statistics 2024-09-18 Abhishek Mandal , Abhisek Chakraborty

Big data presents potential but unresolved value as a source for analysis and inference. However,selection bias, present in many of these datasets, needs to be accounted for so that appropriate inferences can be made on the target…

Methodology · Statistics 2025-01-09 Lyndon Ang , Robert Clark , Bronwyn Loong , Anders Holmberg

We focus on the problem of generalizing a causal effect estimated on a randomized controlled trial (RCT) to a target population described by a set of covariates from observational data. Available methods such as inverse propensity sampling…

Methodology · Statistics 2023-02-27 Imke Mayer , Julie Josse , Traumabase Group

Excess hazard modeling is one of the main tools in population-based cancer survival research. Indeed, this setting allows for direct modeling of the survival due to cancer even in the absence of reliable information on the cause of death,…

Methodology · Statistics 2022-04-12 A. Eletti , G. Marra , M. Quaresma , R. Radice , F. J. Rubio

Statistical integration of diverse data sources is an essential step in the building of generalizable prediction tools, especially in precision health. The invariant features model is a new paradigm for multi-source data integration which…

Methodology · Statistics 2025-03-05 Parker Knight , Ndey Isatou Jobe , Rui Duan

The identification of predictive biomarkers from a large scale of covariates for subgroup analysis has attracted fundamental attention in medical research. In this article, we propose a generalized penalized regression method with a novel…

Methodology · Statistics 2019-04-29 Chong Ma , Wenxuan Deng , Shuangge Ma , Ray Liu , Kevin Galinsky

International comparisons of hierarchical time series data sets based on survey data, such as annual country-level estimates of school enrollment rates, can suffer from large amounts of missing data due to differing coverage of surveys…

Methodology · Statistics 2025-03-31 Daphne H. Liu , Adrian E. Raftery

When multitudes of features can plausibly be associated with a response, both privacy considerations and model parsimony suggest grouping them to increase the predictive power of a regression model. Specifically, the identification of…

Methodology · Statistics 2024-05-07 Brandon Woosuk Park , Anand N. Vidyashankar , Tucker S. McElroy

In cancer research, profiling studies have been extensively conducted, searching for genes/SNPs associated with prognosis. Cancer is a heterogeneous disease. Examining similarity and difference in the genetic basis of multiple subtypes of…

Methodology · Statistics 2013-04-18 Jin Liu , Jian Huang , Yawei Zhang , Qing Lan , Nathaniel Rothman , Tongzhang Zheng , Shuangge Ma

Heterogeneous data are now ubiquitous in many applications in which correctly identifying the subgroups from a heterogeneous population is critical. Although there is an increasing body of literature on subgroup detection, existing methods…

Methodology · Statistics 2025-12-09 Jie Wu , Bo Zhang , Daoji Li , Zemin Zheng

Survival outcomes are common in comparative effectiveness studies and require unique handling because they are usually incompletely observed due to right-censoring. A ``once for all'' approach for causal inference with survival outcomes…

Methodology · Statistics 2021-12-21 Shuxi Zeng , Fan Li , Liangyuan Hu , Fan Li

Understanding the factors that trigger or prevent undesirable health outcomes across patient subpopulations is essential for designing targeted interventions. While randomized controlled trials and expert-led patient interviews are standard…

Artificial Intelligence · Computer Science 2026-05-28 Shishir Adhikari , Guido Muscioni , Mark Shapiro , Plamen Petrov , Elena Zheleva

There has been a lot of work fitting Ising models to multivariate binary data in order to understand the conditional dependency relationships between the variables. However, additional covariates are frequently recorded together with the…

Machine Learning · Statistics 2012-09-28 Jie Cheng , Elizaveta Levina , Pei Wang , Ji Zhu

Estimating heterogeneous treatment effects is an important problem across many domains. In order to accurately estimate such treatment effects, one typically relies on data from observational studies or randomized experiments. Currently,…

Machine Learning · Statistics 2022-02-28 Tobias Hatt , Jeroen Berrevoets , Alicia Curth , Stefan Feuerriegel , Mihaela van der Schaar

Randomized clinical trials typically aim to estimate a marginal treatment effect. While covariate adjustment can improve precision, it may change the estimand in nonlinear models due to noncollapsibility, leading to conditional rather than…

Methodology · Statistics 2026-05-25 Leticia Wuethrich , Torsten Hothorn