English

Scaling Analysis of Affinity Propagation

Artificial Intelligence 2013-05-29 v1 Statistical Mechanics

Abstract

We analyze and exploit some scaling properties of the Affinity Propagation (AP) clustering algorithm proposed by Frey and Dueck (2007). First we observe that a divide and conquer strategy, used on a large data set hierarchically reduces the complexity O(N2){\cal O}(N^2) to O(N(h+2)/(h+1)){\cal O}(N^{(h+2)/(h+1)}), for a data-set of size NN and a depth hh of the hierarchical strategy. For a data-set embedded in a dd-dimensional space, we show that this is obtained without notably damaging the precision except in dimension d=2d=2. In fact, for dd larger than 2 the relative loss in precision scales like N(2d)/(h+1)dN^{(2-d)/(h+1)d}. Finally, under some conditions we observe that there is a value ss^* of the penalty coefficient, a free parameter used to fix the number of clusters, which separates a fragmentation phase (for s<ss<s^*) from a coalescent one (for s>ss>s^*) of the underlying hidden cluster structure. At this precise point holds a self-similarity property which can be exploited by the hierarchical strategy to actually locate its position. From this observation, a strategy based on \AP can be defined to find out how many clusters are present in a given dataset.

Keywords

Cite

@article{arxiv.0910.1800,
  title  = {Scaling Analysis of Affinity Propagation},
  author = {Cyril Furtlehner and Michele Sebag and Xiangliang Zhang},
  journal= {arXiv preprint arXiv:0910.1800},
  year   = {2013}
}

Comments

28 pages, 14 figures, Inria research report

R2 v1 2026-06-21T13:56:26.458Z