中文
相关论文

相关论文: The `Why' behind including `Y' in your imputation …

200 篇论文

Iterative imputation, in which variables are imputed one at a time each given a model predicting from all the others, is a popular technique that can be convenient and flexible, as it replaces a potentially difficult multivariate modeling…

统计理论 · 数学 2012-04-04 Jingchen Liu , Andrew Gelman , Jennifer Hill , Yu-Sung Su

Missing data are ubiquitous in the era of big data and, if inadequately handled, are known to lead to biased findings and have deleterious impact on data-driven decision makings. To mitigate its impact, many missing value imputation methods…

机器学习 · 计算机科学 2021-10-26 Yiliang Zhang , Qi Long

Multivariate bounded discrete data arises in many fields. In the setting of dementia studies, such data is collected when individuals complete neuropsychological tests. We outline a modeling and inference procedure that can model the joint…

统计方法学 · 统计学 2026-02-10 Daniel Suen , Yen-Chi Chen

Learning models that can handle distribution shifts is a key challenge in domain generalization. Invariance learning, an approach that focuses on identifying features invariant across environments, improves model generalization by capturing…

机器学习 · 统计学 2026-05-11 Yiran Jia , Jelena Bradic

In some multivariate problems with missing data, pairs of variables exist that are never observed together. For example, some modern biological tools can produce data of this form. As a result of this structure, the covariance matrix is…

统计方法学 · 统计学 2013-08-13 Max Grazier G'Sell , Shai S. Shen-Orr , Robert Tibshirani

Penalized regression methods, such as lasso and elastic net, are used in many biomedical applications when simultaneous regression coefficient estimation and variable selection is desired. However, missing data complicates the…

Ecological Momentary Assessments (EMA) capture real-time thoughts and behaviors in natural settings, producing rich longitudinal data for statistical and physiological analyses. However, the robustness of these analyses can be compromised…

统计方法学 · 统计学 2023-11-21 Yiheng Wei , Donald Hedeker

Inferring causal effects of a treatment, intervention or policy from observational data is central to many applications. However, state-of-the-art methods for causal inference seldom consider the possibility that covariates have missing…

统计方法学 · 统计学 2020-02-26 Imke Mayer , Julie Josse , Félix Raimundo , Jean-Philippe Vert

Multiple imputation is widely used to handle missing data. Although Rubin's combining rule is simple, it is not clear whether or not the standard multiple imputation inference is consistent when coupled with the commonly-used full sample…

统计方法学 · 统计学 2023-01-03 Qian Guan , Shu Yang

How should researchers analyze randomized experiments in which the main outcome is latent and measured in multiple ways but each measure contains some degree of error? We first identify a critical study-specific noncomparability problem in…

计量经济学 · 经济学 2026-01-13 Jiawei Fu , Donald P. Green

This research deals with the estimation and imputation of missing data in longitudinal models with a Poisson response variable inflated with zeros. A methodology is proposed that is based on the use of maximum likelihood, assuming that data…

统计方法学 · 统计学 2024-09-18 D. S. Martinez-Lobo , O. O. Melo , N. A. Cruz

The problem of missingness in observational data is ubiquitous. When the confounders are missing at random, multiple imputation is commonly used; however, the method requires congeniality conditions for valid inferences, which may not be…

统计方法学 · 统计学 2020-07-10 Nathan Corder , Shu Yang

Multivariate density estimation is a popular technique in statistics with wide applications including regression models allowing for heteroskedasticity in conditional variances. The estimation problems become more challenging when…

统计方法学 · 统计学 2018-08-15 Zhen Li , Lili Wu , Weilian Zhou , Sujit Ghosh

Covariate adjustment can improve precision in analyzing randomized experiments. With fully observed data, regression adjustment and propensity score weighting are asymptotically equivalent in improving efficiency over unadjusted analysis.…

统计方法学 · 统计学 2024-03-06 Anqi Zhao , Peng Ding , Fan Li

Many scientific questions in biomedical, environmental, and psychological research involve understanding the effects of multiple factors on outcomes. While factorial experiments are ideal for this purpose, randomized controlled treatment…

统计方法学 · 统计学 2025-12-03 Ruoqi Yu , Peng Ding

Imputation models sometimes use auxiliary variables that, though not part of the planned analysis, can improve the accuracy of imputed values and the efficiency of point estimates. A recent article, using evidence from simulations, argued…

统计方法学 · 统计学 2013-11-22 Paul von Hippel , Jamie Lynch

There has been a lot of work fitting Ising models to multivariate binary data in order to understand the conditional dependency relationships between the variables. However, additional covariates are frequently recorded together with the…

机器学习 · 统计学 2012-09-28 Jie Cheng , Elizaveta Levina , Pei Wang , Ji Zhu

The interventional effects approach to causal mediation analysis is increasingly common in epidemiologic research, given its potential to address policy-relevant questions about hypothetical mediator interventions. Multiple imputation (MI)…

Many real-world datasets contain missing entries and mixed data types including categorical and ordered (e.g. continuous and ordinal) variables. Imputing the missing entries is necessary, since many data analysis pipelines require complete…

统计方法学 · 统计学 2022-10-14 Yuxuan Zhao , Alex Townsend , Madeleine Udell

BACKGROUND: As databases grow larger, it becomes harder to fully control their collection, and they frequently come with missing values: incomplete observations. These large databases are well suited to train machine-learning models, for…

机器学习 · 计算机科学 2022-02-23 Alexandre Perez-Lebel , Gaël Varoquaux , Marine Le Morvan , Julie Josse , Jean-Baptiste Poline