English
Related papers

Related papers: Multiple imputation of missing covariates when usi…

200 papers

Interval-censoring frequently occurs in studies of chronic diseases where disease status is inferred from intermittently collected biomarkers. Although many methods have been developed to analyze such data, they typically assume perfect…

Methodology · Statistics 2026-05-26 Yuhao Deng , Donglin Zeng , Yuanjia Wang

When using multiple imputation (MI) for missing data, maintaining compatibility between the imputation model and substantive analysis is important for avoiding bias. For example, some causal inference methods incorporate an outcome model…

Propensity score methods are an important tool to help reduce confounding in non-experimental studies. Most propensity score methods assume that covariates are measured without error. However, covariates are often measured with error, which…

Methodology · Statistics 2017-06-08 Hwanhee Hong , David A. Aaby , Juned Siddique , Elizabeth A. Stuart

Population adjustment methods such as matching-adjusted indirect comparison (MAIC) are increasingly used to compare marginal treatment effects when there are cross-trial differences in effect modifiers and limited patient-level data. MAIC…

Methodology · Statistics 2022-05-12 Antonio Remiro-Azócar , Anna Heath , Gianluca Baio

Studies of the effects of medical interventions increasingly take place in distributed research settings using data from multiple clinical data sources including electronic health records and administrative claims. In such settings, privacy…

Methodology · Statistics 2021-01-06 Martijn J. Schuemie , Yong Chen , David Madigan , Marc A. Suchard

Gaussian Mixture models (GMMs) are a powerful tool for clustering, classification and density estimation when clustering structures are embedded in the data. The presence of missing values can largely impact the GMMs estimation process,…

Machine Learning · Statistics 2020-06-05 Alessio Serafini , Thomas Brendan Murphy , Luca Scrucca

We consider survival data in the presence of a cure fraction, meaning that some subjects will never experience the event of interest. We assume a mixture cure model consisting of two sub-models: one for the probability of being uncured…

Methodology · Statistics 2023-03-17 Eni Musta , Tsz Pang Yuen

Analysts are often confronted with censoring, wherein some variables are not observed at their true value, but rather at a value that is known to fall above or below that truth. While much attention has been given to the analysis of…

Methodology · Statistics 2024-05-31 Sarah C. Lotspeich , Kyle F. Grosser , Tanya P. Garcia

Missing data represents a fundamental challenge in machine learning applications, often reducing model performance and reliability. This problem is particularly acute in fields like bioinformatics and clinical machine learning, where…

Machine Learning · Computer Science 2025-09-04 Fatemeh Azad , Zoran Bosnić , Matjaž Kukar

Many modern estimators require bootstrapping to calculate confidence intervals because either no analytic standard error is available or the distribution of the parameter of interest is non-symmetric. It remains however unclear how to…

Methodology · Statistics 2018-09-13 Michael Schomaker , Christian Heumann

Estimation of the mean vector and covariance matrix is of central importance in the analysis of multivariate data. In the framework of generalized linear models, usually the variances are certain functions of the means with the normal…

Methodology · Statistics 2023-01-25 Anupam Kundu , Mohsen Pourahmadi

There has been an increasing interest in using cell and gene therapy (CGT) to treat/cure difficult diseases. The hallmark of CGT trials are the small sample size and extremely high efficacy. Due to the innovation and novelty of such…

Applications · Statistics 2025-10-23 Yaoyuan Vincent Tan , Gang Xu , Chenkun Wang

Missing data frequently occurs in datasets across various domains, such as medicine, sports, and finance. In many cases, to enable proper and reliable analyses of such data, the missing values are often imputed, and it is necessary that the…

In a Cox model, the partial likelihood, as the product of a series of conditional probabilities, is used to estimate the regression coefficients. In practice, those conditional probabilities are approximated by risk score ratios based on a…

Methodology · Statistics 2025-02-27 Youngjin Cho , Yili Hong , Pang Du

Observational longitudinal data on treatments and covariates are increasingly used to investigate treatment effects, but are often subject to time-dependent confounding. Marginal structural models (MSMs), estimated using inverse probability…

Methodology · Statistics 2020-02-11 Ruth H. Keogh , Shaun R. Seaman , Jon Michael Gran , Stijn Vansteelandt

Data collected in clinical trials are often composed of multiple types of variables. For example, laboratory measurements and vital signs are longitudinal data of continuous or categorical variables, adverse events may be recurrent events,…

Methodology · Statistics 2023-01-12 Tuo Wang , Rachel Zilinskas , Ying Li , Yongming Qu

Mixed Probit models are widely applied in many fields where prediction of a binary response is of interest. Typically, the random effects are assumed to be independent but this is seldom the case for many real applications. In the credit…

Applications · Statistics 2019-11-18 Elisa Tosetti , Veronica Vinciotti

Missing data imputation forms the first critical step of many data analysis pipelines. The challenge is greatest for mixed data sets, including real, Boolean, and ordinal data, where standard techniques for imputation fail basic sanity…

Methodology · Statistics 2020-06-17 Yuxuan Zhao , Madeleine Udell

The case-cohort design allows analysis of multiple endpoints and only requires covariates to be measured for cases and non-cases in a random subcohort from the cohort. Stratification of subcohort sampling and weight calibration increase…

Applications · Statistics 2024-02-15 Lola Etievant , Mitchell H. Gail

Although randomized experiments are widely regarded as the gold standard for estimating causal effects, missing data of the pretreatment covariates makes it challenging to estimate the subgroup causal effects. When the missing data…

Statistics Theory · Mathematics 2014-01-08 Peng Ding , Zhi Geng