Extractive summarization is very useful for physicians to better manage and digest Electronic Health Records (EHRs). However, the training of a supervised model requires disease-specific medical background and is thus very expensive. We studied how to utilize the intrinsic correlation between multiple EHRs to generate pseudo-labels and train a supervised model with no external annotation. Experiments on real-patient data validate that our model is effective in summarizing crucial disease-specific information for patients.
@article{arxiv.1811.08040,
title = {Unsupervised Pseudo-Labeling for Extractive Summarization on Electronic Health Records},
author = {Xiangan Liu and Keyang Xu and Pengtao Xie and Eric Xing},
journal= {arXiv preprint arXiv:1811.08040},
year = {2018}
}
Comments
Machine Learning for Health (ML4H) Workshop at NeurIPS 2018 arXiv:1811.07216