English

$k$-POD: A Method for $k$-Means Clustering of Missing Data

Computation 2018-06-07 v3 Methodology

Abstract

The kk-means algorithm is often used in clustering applications but its usage requires a complete data matrix. Missing data, however, is common in many applications. Mainstream approaches to clustering missing data reduce the missing data problem to a complete data formulation through either deletion or imputation but these solutions may incur significant costs. Our kk-POD method presents a simple extension of kk-means clustering for missing data that works even when the missingness mechanism is unknown, when external information is unavailable, and when there is significant missingness in the data.

Keywords

Cite

@article{arxiv.1411.7013,
  title  = {$k$-POD: A Method for $k$-Means Clustering of Missing Data},
  author = {Jocelyn T. Chi and Eric C. Chi and Richard G. Baraniuk},
  journal= {arXiv preprint arXiv:1411.7013},
  year   = {2018}
}

Comments

26 pages, 7 tables

R2 v1 2026-06-22T07:12:15.522Z