Global $k$-means$++$: an effective relaxation of the global $k$-means clustering algorithm
Abstract
The -means algorithm is a prevalent clustering method due to its simplicity, effectiveness, and speed. However, its main disadvantage is its high sensitivity to the initial positions of the cluster centers. The global -means is a deterministic algorithm proposed to tackle the random initialization problem of k-means but its well-known that requires high computational cost. It partitions the data to clusters by solving all -means sub-problems incrementally for all . For each cluster problem, the method executes the -means algorithm times, where is the number of datapoints. In this paper, we propose the \emph{global -means\texttt{++}} clustering algorithm, which is an effective way of acquiring quality clustering solutions akin to those of global -means with a reduced computational load. This is achieved by exploiting the center selection probability that is effectively used in the -means\texttt{++} algorithm. The proposed method has been tested and compared in various benchmark datasets yielding very satisfactory results in terms of clustering quality and execution speed.
Keywords
Cite
@article{arxiv.2211.12271,
title = {Global $k$-means$++$: an effective relaxation of the global $k$-means clustering algorithm},
author = {Georgios Vardakas and Aristidis Likas},
journal= {arXiv preprint arXiv:2211.12271},
year = {2023}
}