English
Related papers

Related papers: Scalable and Efficient Multiple Imputation for Cas…

200 papers

We consider logistic regression including two sets of discrete or categorical covariates that are missing at random (MAR) separately or simultaneously. We examine the asymptotic properties of two multiple imputation (MI) estimators, given…

Methodology · Statistics 2023-12-19 Phuoc-Loc Tran , Shen-Ming Lee , Truong-Nhat Le , Chin-Shang Li

We consider quantile estimation in a semi-supervised setting, characterized by two available data sets: (i) a small or moderate sized labeled data set containing observations for a response and a set of possibly high dimensional covariates,…

Methodology · Statistics 2024-08-15 Abhishek Chakrabortty , Guorong Dai , Raymond J. Carroll

IMPORTANCE: Feature selection with respect to time-to-event outcomes is one of the fundamental problems in clinical trials and biomarker discovery studies. But it's unclear which statistical methods should be used when sample size is small…

Methodology · Statistics 2022-10-17 Rong Lu

Extraordinary amounts of data are being produced in many branches of science. Proven statistical methods are no longer applicable with extraordinary large data sets due to computational limitations. A critical step in big data analysis is…

Methodology · Statistics 2019-06-27 HaiYing Wang , Min Yang , John Stufken

There has been an increasing interest in using cell and gene therapy (CGT) to treat/cure difficult diseases. The hallmark of CGT trials are the small sample size and extremely high efficacy. Due to the innovation and novelty of such…

Applications · Statistics 2025-10-23 Yaoyuan Vincent Tan , Gang Xu , Chenkun Wang

Importance sampling (IS) is a Monte Carlo technique that relies on weighted samples, simulated from a proposal distribution, to estimate intractable integrals. The quality of the estimators improves with the number of samples. However, for…

Computation · Statistics 2022-07-18 Medha Agarwal , Dootika Vats , Víctor Elvira

Large observational datasets, including those derived from electronic health records, are a valuable resource for medical research but are often affected by missingness, measurement error, and misclassification. Two-phase sampling with…

Methodology · Statistics 2026-03-23 Jasper B. Yang , Bryan E. Shepherd , Thomas Lumley , Pamela A. Shaw

This work presents a cost-effective technique for designing robust adaptive beamforming algorithms based on efficient covariance matrix reconstruction with iterative spatial power spectrum (CMR-ISPS). The proposed CMR-ISPS approach…

Machine Learning · Computer Science 2023-09-06 S. Mohammadzadeh , V. H. Nascimento , R. C. de Lamare , O. Kukrer

Propensity score methods are increasingly being used to reduce estimation bias of treatment effects for observational studies. Previous research has shown that propensity score methods consistently estimate the marginal hazard ratio for…

Methodology · Statistics 2019-11-19 Haodi Liang , Cecilia Cotton

Sequential directional importance sampling (SDIS) is an efficient adaptive simulation method for estimating failure probabilities. It expresses the failure probability as the product of a group of integrals that are easy to estimate,…

Methodology · Statistics 2024-10-31 Kai Chenga , Iason Papaioannou , Daniel Straub

This paper studies the case of possibly high-dimensional covariates in the regression discontinuity design (RDD) analysis. In particular, we propose estimation and inference methods for the RDD models with covariate selection which perform…

Econometrics · Economics 2026-01-21 Yoichi Arai , Taisuke Otsu , Myung Hwan Seo

Missing data represents a fundamental challenge in machine learning applications, often reducing model performance and reliability. This problem is particularly acute in fields like bioinformatics and clinical machine learning, where…

Machine Learning · Computer Science 2025-09-04 Fatemeh Azad , Zoran Bosnić , Matjaž Kukar

Selection bias can hinder accurate estimation of association parameters in binary disease risk models using non-probability samples like electronic health records (EHRs). The issue is compounded when participants are recruited from multiple…

Feature selection is a critical task in machine learning and statistics. However, existing feature selection methods either (i) rely on parametric methods such as linear or generalized linear models, (ii) lack theoretical false discovery…

Machine Learning · Statistics 2025-07-18 Omar Melikechi , David B. Dunson , Jeffrey W. Miller

Missing data is a pervasive problem in epidemiology, with multiple imputation (MI) a commonly used analysis method. MI is valid when data are missing at random (MAR). However, definitions of MAR with multiple incomplete variables are not…

Methodology · Statistics 2025-04-14 Paul Madley-Dowd , Rachael A. Hughes , Maya B. Mathur , Jon Heron , Kate Tilling

Clustered multistate process data are commonly encountered in multicenter observational studies and clinical trials. A clinically important estimand with such data is the marginal probability of being in a particular transient state as a…

Methodology · Statistics 2022-09-05 Wenxian Zhou , Giorgos Bakoyannis , Ying Zhang , Constantin T Yiannoutsos

A prevalent feature of high-dimensional data is the dependence among covariates, and model selection is known to be challenging when covariates are highly correlated. To perform model selection for the high-dimensional Cox proportional…

Methodology · Statistics 2022-10-04 Pierre Bayle , Jianqing Fan

Estimating the expectations of functionals applied to sums of random variables (RVs) is a well-known problem encountered in many challenging applications. Generally, closed-form expressions of these quantities are out of reach. A naive…

Information Theory · Computer Science 2022-10-27 Eya Ben Amar , Nadhir Ben Rached , Abdul-Lateef Haji-Ali , Raúl Tempone

An MCMC simulation method based on a two stage delayed rejection Metropolis-Hastings algorithm is proposed to estimate a factor multivariate stochastic volatility model. The first stage uses kstep iteration towards the mode, with k small,…

Computation · Statistics 2010-02-11 Weijun Xu , Li Yang , Robert Kohn

The Cox regression, a semi-parametric method of survival analysis, is extremely popular in biomedical applications. The proportional hazards assumption is a key requirement in the Cox model. To accommodate non-proportional hazards, we…

Methodology · Statistics 2022-06-13 Alexander Begun , Elena Kulinskaya