English

Scalable Dirichlet Process Mixture Models with Unknown Concentration and Adaptive Covariance for High-Dimensional Clustering Applied to Leukemia Transcriptomics

Methodology 2026-02-19 v2

Abstract

We propose a novel method that performs adaptive clustering with DPMM using collapsed VI, while incorporating weakly-informative priors for DP concentration parameter alpha and base distribution G0. We illustrate the importance of G0 covariance structure and prior choice by considering different parameterisations of the data covariance matrix. On high-dimensional Gaussian simulations, our model demonstrates substantially faster convergence than a state-of-the-art MCMC splice sampler. We further evaluate performances on Negative Binomial simulations and conduct sensitivity analyses to assess robustness on realistic data conditions. Application to a publicly available leukemia transcriptomic data set comprising 72 samples and 2,194 gene expression successfully recovers every known sub-type, all while identifying additional gene expression-based sub-clusters with meaningful biological interpretation.

Keywords

Cite

@article{arxiv.2601.21106,
  title  = {Scalable Dirichlet Process Mixture Models with Unknown Concentration and Adaptive Covariance for High-Dimensional Clustering Applied to Leukemia Transcriptomics},
  author = {Annesh Pal and Aguirre Mimoun and Rodolphe Thiébaut and Boris P. Hejblum},
  journal= {arXiv preprint arXiv:2601.21106},
  year   = {2026}
}

Comments

22 pages with 5 figures and 1 table

R2 v1 2026-07-01T09:24:46.730Z