English

Group-regularized matrix factorization for fast and reliable module discovery in pan-omics pan-cancer studies

Applications 2026-08-05 v1 Methodology

Abstract

In pan-omics pan-cancer studies, it is critical to identify latent sources of variation that are shared across particular subsets. This task often requires bidimensionally linked data matrices to be decomposed into a sum of block-sparse, low-rank modules. Existing approaches often rely on pre-specified module numbers, ranks, or post-hoc thresholding and can be sensitive to model specification when the underlying sharing structure is complex. To address these issues, we propose GL-BIDIFAC+, a group-regularized matrix factorization framework for discovering partially shared modules. It requires only an upper bound on the latent dimension and encourages module selection through group regularization with theoretically-motivated tuning parameter selection and local support recovery analysis, providing both scalability and principled guidance for module discovery. It also admits a probabilistic interpretation that enables model-based imputation of missing data. Simulation studies demonstrate accurate module recovery and favorable computational performance relative to existing approaches. We further apply GL-BIDIFAC+ to analyze the Cancer Genome Atlas data, where well-established molecular structure provides interpretable biological references. Our analysis distinguishes broad pan-cancer variation, cancer-specific subtype structure, and variation shared across cancers with related tissue origins or histologic features.

Cite

@article{arxiv.2608.04826,
  title  = {Group-regularized matrix factorization for fast and reliable module discovery in pan-omics pan-cancer studies},
  author = {Jun Young Park and Peter W. MacDonald},
  journal= {arXiv preprint arXiv:2608.04826},
  year   = {2026}
}