中文

回归的不可知样本压缩方案

机器学习 2024-02-06 v2 信息论 math.IT 统计理论 机器学习 统计理论

摘要

我们在具有 p\ell_p 损失的不可知回归设定下,其中 p[1,]p\in [1,\infty],获得了有界样本压缩的首批正向结果。我们构造了一个通用的近似样本压缩方案,用于实值函数类,其规模在 fat-shattering 维数上呈指数增长但与样本大小无关。值得注意的是,对于线性回归,构造了规模与维数呈线性的近似压缩。此外,对于 1\ell_1\ell_\infty 损失,我们甚至能给出规模与维数呈线性的高效精确样本压缩方案。我们进一步表明,对于其他每个 p\ell_p 损失,p(1,)p\in (1,\infty),不存在有界规模的精确不可知压缩方案。这细化并推广了 David、Moran 和 Yehudayoff 关于 2\ell_2 损失的负面结果。我们以提出一般性开放问题作结:对于具有 1\ell_1 损失的不可知回归,是否每个函数类都承认规模等于其伪维数的精确压缩方案?对于 2\ell_2 损失,是否每个函数类都承认规模在 fat-shattering 维数上呈多项式的近似压缩方案?这些问题推广了 Warmuth 关于可实现情形分类的经典样本压缩猜想。

关键词

引用

@article{arxiv.1810.01864,
  title  = {Agnostic Sample Compression Schemes for Regression},
  author = {Idan Attias and Steve Hanneke and Aryeh Kontorovich and Menachem Sadigurschi},
  journal= {arXiv preprint arXiv:1810.01864},
  year   = {2024}
}

备注

New results in this version: (1) Approximate agnostic sample compression scheme for function classes with finite fat-shattering dimension and the $\ell_p$ loss (section 3), (2) Near-optimal approximate compression for linear functions and the $\ell_p$ loss (section 4.1) The results in sections 4.2 and 4.3 appear in the previous version