中文
相关论文

相关论文: Variable selection with missing data in both covar…

200 篇论文

Advancements in data collection techniques and the heterogeneity of data resources can yield high percentages of missing observations on variables, such as block-wise missing data. Under missing-data scenarios, traditional methods such as…

统计方法学 · 统计学 2022-05-17 Wei Lan , Xuerong Chen , Tao Zou , Chih-Ling Tsai

When analyzing data from randomized clinical trials, covariate adjustment can be used to account for chance imbalance in baseline covariates and to increase precision of the treatment effect estimate. A practical barrier to covariate…

统计方法学 · 统计学 2023-07-04 Chia-Rui Chang , Yue Song , Fan Li , Rui Wang

Few problems in statistics are as perplexing as variable selection in the presence of very many redundant covariates. The variable selection problem is most familiar in parametric environments such as the linear model or additive variants…

统计方法学 · 统计学 2021-02-25 Yi Liu , Veronika Ročková , Yuexi Wang

Missing data is a common challenge when analyzing epidemiological data, and imputation is often used to address this issue. Here, we investigate the scenario where a covariate used in an analysis has missingness and will be imputed. There…

统计方法学 · 统计学 2024-03-04 Lucy D'Agostino McGowan , Sarah C. Lotspeich , Staci A. Hepler

Methods for estimating heterogeneous treatment effect in observational data have largely focused on continuous or binary outcomes, and have been relatively less vetted with survival outcomes. Using flexible machine learning methods in the…

应用统计 · 统计学 2021-07-09 Liangyuan Hu , Jiayi Ji , Fan Li

The standard quantile regression model assumes a linear relationship at the quantile of interest and that all variables are observed. We relax these assumptions by considering a partial linear model while allowing for missing linear…

统计方法学 · 统计学 2016-06-07 Ben Sherwood

A key challenge in estimating causal effects from observational data is handling confounding and is commonly achieved through weighting methods that balance distribution of covariates between treatment and control groups. Weighting…

统计方法学 · 统计学 2025-12-23 Simion De , Jared D. Huling

Measurement error is prevalent across all domains of scientific research where only imprecise observations, rather than the true underlying values, can be obtained. For example, estimates of human microbiome diversity are based on small…

统计方法学 · 统计学 2026-03-10 Kevin McCoy , Zachary Wooten , Christine B. Peterson

The use of flexible machine-learning (ML) models to generate imputations of missing data within the framework of Multiple Imputation (MI) has recently gained traction, particularly in observational settings. For randomised controlled trials…

统计方法学 · 统计学 2025-10-07 Mia S. Tackney , Jonathan W. Bartlett , Elizabeth Williamson , Kim May Lee

In many life science experiments or medical studies, subjects are repeatedly observed and measurements are collected in factorial designs with multivariate data. The analysis of such multivariate data is typically based on multivariate…

统计方法学 · 统计学 2023-05-24 Lubna Amro , Frank Konietschke , Markus Pauly

Instrumental variable (IV) methods are widely used to infer treatment effects in the presence of unmeasured confounding. In this paper, we study nonparametric inference with an IV under a separable binary treatment choice model, which…

统计方法学 · 统计学 2026-02-03 Chan Park , Eric Tchetgen Tchetgen

We investigate the fairness concerns of training a machine learning model using data with missing values. Even though there are a number of fairness intervention methods in the literature, most of them require a complete training set as…

机器学习 · 计算机科学 2022-04-15 Haewon Jeong , Hao Wang , Flavio P. Calmon

Imputing missing values is an important preprocessing step in data analysis, but the literature offers little guidance on how to choose between different imputation models. This letter suggests adopting the imputation model that generates a…

统计方法学 · 统计学 2021-07-13 Moritz Marbach

There is a dearth of robust methods to estimate the causal effects of multiple treatments when the outcome is binary. This paper uses two unique sets of simulations to propose and evaluate the use of Bayesian Additive Regression Trees…

统计方法学 · 统计学 2020-01-22 Liangyuan Hu , Chenyang Gu , Michael Lopez , Jiayi Ji , Juan Wisnivesky

In recent years, theoretical results and simulation evidence have shown Bayesian additive regression trees to be a highly-effective method for nonparametric regression. Motivated by cost-effectiveness analyses in health economics, where…

While achieving high prediction accuracy is a fundamental goal in machine learning, an equally important task is finding a small number of features with high explanatory power. One popular selection technique is permutation importance,…

机器学习 · 统计学 2024-10-02 Min Lu , Hemant Ishwaran

Presence of missing values in a dataset can adversely affect the performance of a classifier. Single and Multiple Imputation are normally performed to fill in the missing values. In this paper, we present several variants of combining…

机器学习 · 计算机科学 2019-10-16 Shehroz S. Khan , Amir Ahmad , Alex Mihailidis

Variable selection for recovering sparsity in nonadditive nonparametric models has been challenging. This problem becomes even more difficult due to complications in modeling unknown interaction terms among high dimensional variables. There…

统计方法学 · 统计学 2012-06-14 Zaili Fang , Inyoung Kim , Patrick Schaumont

We aim to incorporate variable selection routines into variable-by-variable (or sequential) imputation in clustered data to achieve computational improvement in applications with large-scale health data. Specifically, we utilize variable…

统计方法学 · 统计学 2025-04-08 Qiushuang Li , Recai Yucel

The selection of essential variables in logistic regression is vital because of its extensive use in medical studies, finance, economics and related fields. In this paper, we explore four main typologies (test-based, penalty-based,…

统计方法学 · 统计学 2022-05-17 Souvik Bag , Kapil Gupta , Soudeep Deb