English

$k$-Means Clustering for Persistent Homology

Applications 2023-11-28 v4 Optimization and Control Machine Learning

Abstract

Persistent homology is a methodology central to topological data analysis that extracts and summarizes the topological features within a dataset as a persistence diagram; it has recently gained much popularity from its myriad successful applications to many domains. However, its algebraic construction induces a metric space of persistence diagrams with a highly complex geometry. In this paper, we prove convergence of the kk-means clustering algorithm on persistence diagram space and establish theoretical properties of the solution to the optimization problem in the Karush--Kuhn--Tucker framework. Additionally, we perform numerical experiments on various representations of persistent homology, including embeddings of persistence diagrams as well as diagrams themselves and their generalizations as persistence measures; we find that kk-means clustering performance directly on persistence diagrams and measures outperform their vectorized representations.

Keywords

Cite

@article{arxiv.2210.10003,
  title  = {$k$-Means Clustering for Persistent Homology},
  author = {Yueqi Cao and Prudence Leung and Anthea Monod},
  journal= {arXiv preprint arXiv:2210.10003},
  year   = {2023}
}

Comments

21 pages, 6 figures

R2 v1 2026-06-28T03:56:01.308Z