English

Kernel K-means clustering of distributional data

Machine Learning 2025-09-23 v1 Machine Learning Computation

Abstract

We consider the problem of clustering a sample of probability distributions from a random distribution on Rp\mathbb R^p. Our proposed partitioning method makes use of a symmetric, positive-definite kernel kk and its associated reproducing kernel Hilbert space (RKHS) H\mathcal H. By mapping each distribution to its corresponding kernel mean embedding in H\mathcal H, we obtain a sample in this RKHS where we carry out the KK-means clustering procedure, which provides an unsupervised classification of the original sample. The procedure is simple and computationally feasible even for dimension p>1p>1. The simulation studies provide insight into the choice of the kernel and its tuning parameter. The performance of the proposed clustering procedure is illustrated on a collection of Synthetic Aperture Radar (SAR) images.

Keywords

Cite

@article{arxiv.2509.18037,
  title  = {Kernel K-means clustering of distributional data},
  author = {Amparo Baíllo and Jose R. Berrendero and Martín Sánchez-Signorini},
  journal= {arXiv preprint arXiv:2509.18037},
  year   = {2025}
}
R2 v1 2026-07-01T05:50:09.292Z