English

On the Sparsifiability of Correlation Clustering: Approximation Guarantees under Edge Sampling

Machine Learning 2026-02-17 v1 Artificial Intelligence

Abstract

Correlation Clustering (CC) is a fundamental unsupervised learning primitive whose strongest LP-based approximation guarantees require Θ(n3)\Theta(n^3) triangle inequality constraints and are prohibitive at scale. We initiate the study of \emph{sparsification--approximation trade-offs} for CC, asking how much edge information is needed to retain LP-based guarantees. We establish a structural dichotomy between pseudometric and general weighted instances. On the positive side, we prove that the VC dimension of the clustering disagreement class is exactly n1n{-}1, yielding additive ε\varepsilon-coresets of optimal size O~(n/ε2)\tilde{O}(n/\varepsilon^2); that at most (n2)\binom{n}{2} triangle inequalities are active at any LP vertex, enabling an exact cutting-plane solver; and that a sparsified variant of LP-PIVOT, which imputes missing LP marginals via triangle inequalities, achieves a robust 103\frac{10}{3}-approximation (up to an additive term controlled by an empirically computable imputation-quality statistic Γw\overline{\Gamma}_w) once Θ~(n3/2)\tilde{\Theta}(n^{3/2}) edges are observed, a threshold we prove is sharp. On the negative side, we show via Yao's minimax principle that without pseudometric structure, any algorithm observing o(n)o(n) uniformly random edges incurs an unbounded approximation ratio, demonstrating that the pseudometric condition governs not only tractability but also the robustness of CC to incomplete information.

Keywords

Cite

@article{arxiv.2602.13684,
  title  = {On the Sparsifiability of Correlation Clustering: Approximation Guarantees under Edge Sampling},
  author = {Ibne Farabi Shihab and Sanjeda Akter and Anuj Sharma},
  journal= {arXiv preprint arXiv:2602.13684},
  year   = {2026}
}
R2 v1 2026-07-01T10:36:41.542Z