English
Related papers

Related papers: Outcome identification in electronic health record…

200 papers

Clustering is a widely deployed unsupervised learning tool. Model-based clustering is a flexible framework to tackle data heterogeneity when the clusters have different shapes. Likelihood-based inference for mixture distributions often…

Machine Learning · Statistics 2023-05-30 Yubo Zhuang , Xiaohui Chen , Yun Yang

We develop a Bayesian framework for tackling the supervised clustering problem, the generic problem encountered in tasks such as reference matching, coreference resolution, identity uncertainty and record linkage. Our clustering model is…

Machine Learning · Computer Science 2009-07-07 Hal Daumé , Daniel Marcu

Learning from electronic medical records (EMR) is challenging due to their relational nature and the uncertain dependence between a patient's past and future health status. Statistical relational learning is a natural fit for analyzing EMRs…

Machine Learning · Computer Science 2012-07-03 Jesse Davis , Vitor Santos Costa , Peggy Peissig , Michael Caldwell , Elizabeth Berg , David Page

In electronic health records (EHRs), latent subgroups of patients may exhibit distinctive patterning in their longitudinal health trajectories. For such data, growth mixture models (GMMs) enable classifying patients into different latent…

Methodology · Statistics 2022-01-12 Rebecca Anthopolos , Ying Wei , Qixuan Chen

In longitudinal studies using routinely collected data, such as electronic health records (EHRs), patients tend to have more measurements when they are unwell; this informative observation pattern may lead to bias. While semi-parametric…

Methodology · Statistics 2024-10-02 Rose Garrett , Brian Feldman , Eleanor Pullenayegum

We develop a dependent Dirichlet process (DDP) model for repeated measures multiple membership (MM) data. This data structure arises in studies under which an intervention is delivered to each client through a sequence of elements which…

Applications · Statistics 2013-12-09 Terrance D. Savitsky , Susan M. Paddock

We present a sequential Monte Carlo sampler for coalescent based Bayesian hierarchical clustering. The model is appropriate for multivariate non-\iid data and our approach offers a substantial reduction in computational cost when compared…

Applications · Statistics 2016-08-11 Ricardo Henao , Joseph E. Lucas

In many supervised learning applications, the response consists of both continuous and binary outcomes. Studies have shown that jointly modeling such mixed-type responses can substantially improve predictive performance compared to separate…

Methodology · Statistics 2026-03-13 Yu Wang , Ran Jin , Lulu Kang

This paper addresses patient heterogeneity associated with prediction problems in biomedical applications. We propose a systematic hypothesis testing approach to determine the existence of patient subgroup structure and the number of…

Methodology · Statistics 2021-01-08 Xu Gao , Weining Shen , Jing Ning , Ziding Feng , Jianhua Hu

During the semiconductor manufacturing process, predicting the yield of the semiconductor is an important problem. Early detection of defective product production in the manufacturing process can save huge production cost. The data…

Applications · Statistics 2021-08-03 Sewon Park , Kyeongwon Lee , Da-Eun Jeong , Heung-Kook Ko , Jaeyong Lee

Chronic kidney disease (CKD) is a gradual loss of renal function over time, and it increases the risk of mortality, decreased quality of life, as well as serious complications. The prevalence of CKD has been increasing in the last couple of…

Machine Learning · Computer Science 2020-11-18 Jinghe Zhang , Kamran Kowsari , Mehdi Boukhechba , James Harrison , Jennifer Lobo , Laura Barnes

Traditional Evidence Deep Learning (EDL) methods rely on static hyperparameter for uncertainty calibration, limiting their adaptability in dynamic data distributions, which results in poor calibration and generalization in high-risk…

Machine Learning · Computer Science 2025-10-13 Zhen Yang , Yansong Ma , Lei Chen

The adoption of digital systems in healthcare has resulted in the accumulation of vast electronic health records (EHRs), offering valuable data for machine learning methods to predict patient health outcomes. However, single-visit records…

Machine Learning · Computer Science 2024-11-19 Eric Yang , Pengfei Hu , Xiaoxue Han , Yue Ning

The development of electronic health records (EHR) systems has enabled the collection of a vast amount of digitized patient data. However, utilizing EHR data for predictive modeling presents several challenges due to its unique…

Machine Learning · Computer Science 2024-08-14 Jiaqi Wang , Junyu Luo , Muchao Ye , Xiaochen Wang , Yuan Zhong , Aofei Chang , Guanjie Huang , Ziyi Yin , Cao Xiao , Jimeng Sun , Fenglong Ma

This study introduces a novel approach for early Type 2 Diabetes Mellitus (T2DM) risk prediction using a tabular transformer (TabTrans) architecture to analyze longitudinal patient data. By processing patients` longitudinal health records…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Sulaiman Khan , Md. Rafiul Biswas , Zubair Shah

Effective management of Type 1 Diabetes requires continuous glucose monitoring and precise insulin adjustments to prevent hyperglycemia and hypoglycemia. With the growing adoption of wearable glucose monitors and mobile health applications,…

Machine Learning · Computer Science 2026-01-22 Giorgia Rigamonti , Mirko Paolo Barbato , Davide Marelli , Paolo Napoletano

Data augmentation by mixing samples, such as Mixup, has widely been used typically for classification tasks. However, this strategy is not always effective due to the gap between augmented samples for training and original samples for…

Machine Learning · Computer Science 2019-06-21 Takuya Shimada , Shoichiro Yamaguchi , Kohei Hayashi , Sosuke Kobayashi

In this paper, we showed that the no-arbitrage condition holds if the market follows the mixture of the geometric Brownian motion (GBM). The mixture of GBM can incorporate heavy-tail behavior of the market. It automatically leads us to…

Methodology · Statistics 2018-05-02 Sourish Das , Aritra Halder , Ananya Lahiri , Dipak K Dey

The paper presents a new perspective on the mixture of Dirichlet process model which allows the recovery of full and correct uncertainty quantification associated with the full model, even after having integrated out the random distribution…

Computation · Statistics 2022-06-20 Blake Moya , Stephen G. Walker

Ensembles of models often yield improvements in system performance. These ensemble approaches have also been empirically shown to yield robust measures of uncertainty, and are capable of distinguishing between different \emph{forms} of…

Machine Learning · Statistics 2019-11-27 Andrey Malinin , Bruno Mlodozeniec , Mark Gales