Databases of electronic health records (EHRs) are increasingly used to inform clinical decisions. Machine learning methods can find patterns in EHRs that are predictive of future adverse outcomes. However, statistical models may be built upon patterns of health-seeking behavior that vary across patient subpopulations, leading to poor predictive performance when training on one patient population and predicting on another. This note proposes two tests to better measure and understand model generalization. We use these tests to compare models derived from two data sources: (i) historical medical records, and (ii) electrocardiogram (EKG) waveforms. In a predictive task, we show that EKG-based models can be more stable than EHR-based models across different patient populations.
@article{arxiv.1812.00210,
title = {Measuring the Stability of EHR- and EKG-based Predictive Models},
author = {Andrew C. Miller and Ziad Obermeyer and Sendhil Mullainathan},
journal= {arXiv preprint arXiv:1812.00210},
year = {2018}
}
Comments
Machine Learning for Health (ML4H) Workshop at NeurIPS 2018 arXiv:cs/0101200