Fast $k$-means Seeding Under The Manifold Hypothesis
Abstract
We study beyond worst case analysis for the -means problem where the goal is to model typical instances of -means arising in practice. Existing theoretical approaches provide guarantees under certain assumptions on the optimal solutions to -means, making them difficult to validate in practice. We propose the manifold hypothesis, where data obtained in ambient dimension concentrates around a low dimensional manifold of intrinsic dimension , as a reasonable assumption to model real world clustering instances. We identify key geometric properties of datasets which have theoretically predictable scaling laws depending on the quantization exponent using techniques from optimum quantization theory. We show how to exploit these regularities to design a fast seeding method called which provides approximate solutions to the -means problem in time ; where the exponent for an input parameter . This allows us to obtain new runtime - quality tradeoffs. We perform a large scale empirical study across various domains to validate our theoretical predictions and algorithm performance to bridge theory and practice for beyond worst case data clustering.
Cite
@article{arxiv.2602.01104,
title = {Fast $k$-means Seeding Under The Manifold Hypothesis},
author = {Poojan Shah and Shashwat Agrawal and Ragesh Jaiswal},
journal= {arXiv preprint arXiv:2602.01104},
year = {2026}
}