English

Random Models for Fuzzy Clustering Similarity Measures

Machine Learning 2025-02-17 v1 Machine Learning

Abstract

The Adjusted Rand Index (ARI) is a widely used method for comparing hard clusterings, but requires a choice of random model that is often left implicit. Several recent works have extended the Rand Index to fuzzy clusterings, but the assumptions of the most common random model is difficult to justify in fuzzy settings. We propose a single framework for computing the ARI with three random models that are intuitive and explainable for both hard and fuzzy clusterings, along with the benefit of lower computational complexity. The theory and assumptions of the proposed models are contrasted with the existing permutation model. Computations on synthetic and benchmark data show that each model has distinct behaviour, meaning that accurate model selection is important for the reliability of results.

Keywords

Cite

@article{arxiv.2312.10270,
  title  = {Random Models for Fuzzy Clustering Similarity Measures},
  author = {Ryan DeWolfe and Jeffery L. Andrews},
  journal= {arXiv preprint arXiv:2312.10270},
  year   = {2025}
}
R2 v1 2026-06-28T13:53:14.636Z