中文

Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure

机器学习 2026-08-13 v1 机器学习

摘要

The per-instance Jaccard score, or intersection over union (IoU), is standard in multi-label classification and binary segmentation. With ss labels, its loss matrix has 2s2^s outcomes and reports. Under the convention Jac(,)=1\mathrm{Jac}(\varnothing,\varnothing)=1, we prove that the Jaccard score, shifted-loss, and ordinary loss matrices are nonsingular and that the loss columns have affine dimension 2s12^s-1. The proof combines a finite MinHash Gram representation with Boolean M\"obius inversion. For exact calibration, we prove 2s1CCdim(LJac)2s12^{s-1} \leq \mathrm{CCdim}(L^{\mathrm{Jac}}) \leq 2^s-1. The lower bound uses a factorially weighted distribution with 2s1+12^{s-1}+1 supported outcomes and Bayes-optimal reports. Consequently, every exactly calibrated convex surrogate requires exponentially many prediction coordinates. We also give two polynomial-dimensional approximation guarantees with explicit regret transfers. A new F1F_1-to-Jaccard transfer turns an existing (s2+1)(s^2+1)-dimensional F1F_1 surrogate into a polynomial-time rule with asymptotic Jaccard regret at most 3223-2\sqrt{2}. For any α>0\alpha>0 and 0<ρ<10<\rho<1, a MinHash square-loss surrogate attains Jaccard-regret floor α\alpha uniformly over arbitrary conditional label distributions. With probability at least 1ρ1-\rho, the direct construction has dimension O((s2+slog(1/ρ))/α2)O((s^2+s\log(1/\rho))/\alpha^2), while a signed variant has dimension O((s+log(1/ρ))/α2)O((s+\log(1/\rho))/\alpha^2). Thus zero-regret calibration requires exponential dimension, whereas every fixed additive regret tolerance admits polynomial prediction dimension.

引用

@article{arxiv.2608.13549,
  title  = {Exponential Convex Calibration Dimension for the Multi-Label Jaccard Measure},
  author = {Mingyuan Zhang},
  journal= {arXiv preprint arXiv:2608.13549},
  year   = {2026}
}