English

Clustering, Coding, and the Concept of Similarity

Machine Learning 2024-05-14 v2

Abstract

This paper develops a theory of clustering and coding which combines a geometric model with a probabilistic model in a principled way. The geometric model is a Riemannian manifold with a Riemannian metric, gij(x){g}_{ij}({\bf x}), which we interpret as a measure of dissimilarity. The probabilistic model consists of a stochastic process with an invariant probability measure which matches the density of the sample input data. The link between the two models is a potential function, U(x)U({\bf x}), and its gradient, U(x)\nabla U({\bf x}). We use the gradient to define the dissimilarity metric, which guarantees that our measure of dissimilarity will depend on the probability measure. Finally, we use the dissimilarity metric to define a coordinate system on the embedded Riemannian manifold, which gives us a low-dimensional encoding of our original data.

Keywords

Cite

@article{arxiv.1401.2411,
  title  = {Clustering, Coding, and the Concept of Similarity},
  author = {L. Thorne McCarty},
  journal= {arXiv preprint arXiv:1401.2411},
  year   = {2024}
}

Comments

Revised and expanded in response to referee reports. Current version: 65 pages, 18 figures

R2 v1 2026-06-22T02:43:03.364Z