English

Efficient model-based clustering with coalescents: Application to multiple outcomes using medical records data

Applications 2016-08-11 v1

Abstract

We present a sequential Monte Carlo sampler for coalescent based Bayesian hierarchical clustering. The model is appropriate for multivariate non-\iid data and our approach offers a substantial reduction in computational cost when compared to the original sampler. We also propose a quadratic complexity approximation that in practice shows almost no loss in performance compared to its counterpart. Our formulation leads to a greedy algorithm that exhibits performance improvement over other greedy algorithms, particularly in small data sets. We incorporate the Coalescent into a hierarchical regression model that allows joint modeling of multiple correlated outcomes. The approach does not require {\em a priori} knowledge of either the degree or structure of the correlation and, as a byproduct, generates additional models for a subset of the composite outcomes. We demonstrate the utility of the approach by predicting multiple different types of outcomes using medical records data from a cohort of diabetic patients.

Keywords

Cite

@article{arxiv.1608.03191,
  title  = {Efficient model-based clustering with coalescents: Application to multiple outcomes using medical records data},
  author = {Ricardo Henao and Joseph E. Lucas},
  journal= {arXiv preprint arXiv:1608.03191},
  year   = {2016}
}

Comments

arXiv admin note: substantial text overlap with arXiv:1204.4708

R2 v1 2026-06-22T15:16:55.466Z