一种比较多重插补技术的方法:基于美国国家级 COVID 队列协作平台的案例研究
人工智能
2022-09-27 v2 计算机与社会
应用统计
摘要
从电子健康记录获取的医疗数据集已被证明在评估患者预测因子与关注结局之间的关联方面极为有用。然而,这些数据集常在很高比例的病例中存在缺失值,而简单删除这些病例可能引入严重偏倚。出于这些原因,已有多种多重插补算法被提出以试图恢复缺失信息。每种算法各有优劣,目前对于在给定场景下哪种多重插补算法效果最佳尚无共识。此外,各算法参数的选择以及与数据相关的建模选择同样关键且具挑战性。本文提出一种新颖框架,用于在统计分析背景下数值评估处理缺失数据的策略,特别关注多重插补技术。我们在国家 COVID 队列协作平台(N3C)Enclave 提供的 2 型糖尿病患者大型队列上验证了该方法的可行性,探讨了各类患者特征对 COVID-19 相关结局的影响。我们的分析包含经典多重插补技术以及简单的完整病例逆概率加权模型。本文实验表明,我们的方法能有效凸显针对该案例研究最有效且性能最佳的数据缺失处理策略。此外,我们的方法使我们理解了不同模型的行为及其随参数修改的变化方式。我们的方法具有通用性,可应用于不同研究领域及含异构类型的数据集。
引用
@article{arxiv.2206.06444,
title = {A method for comparing multiple imputation techniques: a case study on the U.S. National COVID Cohort Collaborative},
author = {Elena Casiraghi and Rachel Wong and Margaret Hall and Ben Coleman and Marco Notaro and Michael D. Evans and Jena S. Tronieri and Hannah Blau and Bryan Laraway and Tiffany J. Callahan and Lauren E. Chan and Carolyn T. Bramante and John B. Buse and Richard A. Moffitt and Til Sturmer and Steven G. Johnson and Yu Raymond Shao and Justin Reese and Peter N. Robinson and Alberto Paccanaro and Giorgio Valentini and Jared D. Huling and Kenneth Wilkins and : and Tell Bennet and Christopher Chute and Peter DeWitt and Kenneth Gersing and Andrew Girvin and Melissa Haendel and Jeremy Harper and Janos Hajagos and Stephanie Hong and Emily Pfaff and Jane Reusch and Corneliu Antoniescu and Kimberly Robaski},
journal= {arXiv preprint arXiv:2206.06444},
year = {2022}
}