English
Related papers

Related papers: Semiparametric Marginal Regression for Clustered C…

200 papers

Handling incomplete and heterogeneous data remains a central challenge in real-world machine learning, where missing values may follow complex mechanisms (MCAR, MAR, MNAR) and features can be of mixed types (numerical and categorical).…

Machine Learning · Computer Science 2025-07-30 Youran Zhou , Mohamed Reda Bouadjenek , Jonathan Wells , Sunil Aryal

Single-index models or time-to-event models are frequently applied in empirical research. These models are non-identifiable in presence of unknown (dependent) censoring or competing risks and do not give informative results in empirical…

Methodology · Statistics 2026-03-25 Jia-Han Shih , Simon M. S. Lo , Ralf A. Wilke

Background: Existing guidelines for handling missing data are generally not consistent with the goals of prediction modelling, where missing data can occur at any stage of the model pipeline. Multiple imputation (MI), often heralded as the…

Methodology · Statistics 2022-06-27 Rose Sisk , Matthew Sperrin , Niels Peek , Maarten van Smeden , Glen P. Martin

In fitting a mixture of linear regression models, normal assumption is traditionally used to model the error and then regression parameters are estimated by the maximum likelihood estimators (MLE). This procedure is not valid if the normal…

Methodology · Statistics 2018-11-06 Yanyuan Ma , Shaoli Wang , Lin Xu , Weixin Yao

Marginal structural models (MSMs) are commonly used to estimate causal intervention effects in longitudinal non-randomised studies. A common issue when analysing data from observational studies is the presence of incomplete confounder data,…

Methodology · Statistics 2019-12-02 Clemence Leyrat , James R Carpenter , Sebastien Bailly , Elizabeth J Willamson

We consider the complex data modeling problem motivated by the zero-inflated and overdispersed data from microbiome studies. Analyzing how microbiome abundance is associated with human biological features, such as BMI, is of great…

Methodology · Statistics 2025-03-31 Zirui Wang , Tianying Wang

Frequently, empirical studies are plagued with missing data. When the data are missing not at random, the parameter of interest is not identifiable in general. Without additional assumptions, we can derive bounds of the parameters of…

Methodology · Statistics 2018-09-12 Zhichao Jiang , Peng Ding

Semi-supervised learning has received increasingly attention in statistics and machine learning. In semi-supervised learning settings, a labeled data set with both outcomes and covariates and an unlabeled data set with covariates only are…

Machine Learning · Statistics 2024-02-26 Zhuojun Quan , Yuanyuan Lin , Kani Chen , Wen Yu

Missing data is a common problem in medical research, and is commonly addressed using multiple imputation. Although traditional imputation methods allow for valid statistical inference when data are missing at random (MAR), their…

Classical semiparametric inference with missing outcome data is not robust to contamination of the observed data and a single observation can have arbitrarily large influence on estimation of a parameter of interest. This sensitivity is…

Methodology · Statistics 2021-03-02 Eva Cantoni , Xavier de Luna

For time-to-event data with finitely many competing risks, the proportional hazards model has been a popular tool for relating the cause-specific outcomes to covariates [Prentice et al. Biometrics 34 (1978) 541--554]. This article studies…

Statistics Theory · Mathematics 2009-03-04 Yanqing Sun , Peter B. Gilbert , Ian W. McKeague

Semi-competing risks data arise when both non-terminal and terminal events are considered in a model. Such data with multiple events of interest are frequently encountered in medical research and clinical trials. In this framework, terminal…

Methodology · Statistics 2022-11-21 Fatemeh Mahmoudi , Xuewen Lu

Recently, there has been significant interest in linear regression in the situation where predictors and responses are not observed in matching pairs corresponding to the same statistical unit as a consequence of separate data collection…

Methodology · Statistics 2019-10-04 Martin Slawski , Guoqing Diao , Emanuel Ben-David

Semiparametric models are often considered for analyzing longitudinal data for a good balance between flexibility and parsimony. In this paper, we study a class of marginal partially linear quantile models with possibly varying…

Statistics Theory · Mathematics 2009-11-19 Huixia Judy Wang , Zhongyi Zhu , Jianhui Zhou

Model averaging has demonstrated superior performance for ensemble forecasting in high-dimensional framework, its extension to incomplete datasets remains a critical but underexplored challenge. Moreover, identifying the parsimonious model…

Methodology · Statistics 2025-09-03 Wei Xiong , Dianliang Deng , Dehui Wang

Clinical prediction models must be developed using sufficiently large datasets to minimise overfitting and ensure robust predictive performance. Existing sample size calculations assume complete predictor data for all included participants,…

We consider generalized linear regression analysis with left-censored covariate due to the lower limit of detection. Complete case analysis by eliminating observations with values below limit of detection yields valid estimates for…

Methodology · Statistics 2014-12-09 Shengchun Kong , Bin Nan

Suppose we are interested in the mean of an outcome that is subject to nonignorable nonresponse. This paper develops new semiparametric estimation methods with instrumental variables which affect nonresponse, but not the outcome. The…

Methodology · Statistics 2024-08-20 Baoluo Sun , Wang Miao , Deshanee S. Wickramarachchi

We study semiparametric efficiency bounds and efficient estimation of parameters defined through general moment restrictions with missing data. Identification relies on auxiliary data containing information about the distribution of the…

Statistics Theory · Mathematics 2008-04-04 Xiaohong Chen , Han Hong , Alessandro Tarozzi

In experimental and observational data settings, researchers often have limited knowledge of the reasons for missing outcomes. To address this uncertainty, we propose bounds on causal effects for missing outcomes, accommodating the scenario…

Methodology · Statistics 2026-03-19 Max Rubinstein , Denis Agniel , Larry Han , Marcela Horvitz-Lennon , Sharon-Lise Normand