English

Improved Approximations for Euclidean $k$-means and $k$-median, via Nested Quasi-Independent Sets

Data Structures and Algorithms 2022-04-13 v2 Computational Geometry Machine Learning

Abstract

Motivated by data analysis and machine learning applications, we consider the popular high-dimensional Euclidean kk-median and kk-means problems. We propose a new primal-dual algorithm, inspired by the classic algorithm of Jain and Vazirani and the recent algorithm of Ahmadian, Norouzi-Fard, Svensson, and Ward. Our algorithm achieves an approximation ratio of 2.4062.406 and 5.9125.912 for Euclidean kk-median and kk-means, respectively, improving upon the 2.633 approximation ratio of Ahmadian et al. and the 6.1291 approximation ratio of Grandoni, Ostrovsky, Rabani, Schulman, and Venkat. Our techniques involve a much stronger exploitation of the Euclidean metric than previous work on Euclidean clustering. In addition, we introduce a new method of removing excess centers using a variant of independent sets over graphs that we dub a "nested quasi-independent set". In turn, this technique may be of interest for other optimization problems in Euclidean and p\ell_p metric spaces.

Keywords

Cite

@article{arxiv.2204.04828,
  title  = {Improved Approximations for Euclidean $k$-means and $k$-median, via Nested Quasi-Independent Sets},
  author = {Vincent Cohen-Addad and Hossein Esfandiari and Vahab Mirrokni and Shyam Narayanan},
  journal= {arXiv preprint arXiv:2204.04828},
  year   = {2022}
}

Comments

74 pages. To appear in Symposium on Theory of Computing (STOC), 2022