English

$k$-Variance: A Clustered Notion of Variance

Statistics Theory 2020-12-15 v1 Machine Learning Numerical Analysis Numerical Analysis Machine Learning Statistics Theory

Abstract

We introduce kk-variance, a generalization of variance built on the machinery of random bipartite matchings. KK-variance measures the expected cost of matching two sets of kk samples from a distribution to each other, capturing local rather than global information about a measure as kk increases; it is easily approximated stochastically using sampling and linear programming. In addition to defining kk-variance and proving its basic properties, we provide in-depth analysis of this quantity in several key cases, including one-dimensional measures, clustered measures, and measures concentrated on low-dimensional subsets of Rn\mathbb R^n. We conclude with experiments and open problems motivated by this new way to summarize distributional shape.

Keywords

Cite

@article{arxiv.2012.06958,
  title  = {$k$-Variance: A Clustered Notion of Variance},
  author = {Justin Solomon and Kristjan Greenewald and Haikady N. Nagaraja},
  journal= {arXiv preprint arXiv:2012.06958},
  year   = {2020}
}
R2 v1 2026-06-23T20:55:39.234Z