回归的不可知样本压缩方案
机器学习
2024-02-06 v2 信息论
math.IT
统计理论
机器学习
统计理论
摘要
我们在具有 损失的不可知回归设定下,其中 ,获得了有界样本压缩的首批正向结果。我们构造了一个通用的近似样本压缩方案,用于实值函数类,其规模在 fat-shattering 维数上呈指数增长但与样本大小无关。值得注意的是,对于线性回归,构造了规模与维数呈线性的近似压缩。此外,对于 与 损失,我们甚至能给出规模与维数呈线性的高效精确样本压缩方案。我们进一步表明,对于其他每个 损失,,不存在有界规模的精确不可知压缩方案。这细化并推广了 David、Moran 和 Yehudayoff 关于 损失的负面结果。我们以提出一般性开放问题作结:对于具有 损失的不可知回归,是否每个函数类都承认规模等于其伪维数的精确压缩方案?对于 损失,是否每个函数类都承认规模在 fat-shattering 维数上呈多项式的近似压缩方案?这些问题推广了 Warmuth 关于可实现情形分类的经典样本压缩猜想。
引用
@article{arxiv.1810.01864,
title = {Agnostic Sample Compression Schemes for Regression},
author = {Idan Attias and Steve Hanneke and Aryeh Kontorovich and Menachem Sadigurschi},
journal= {arXiv preprint arXiv:1810.01864},
year = {2024}
}
备注
New results in this version: (1) Approximate agnostic sample compression scheme for function classes with finite fat-shattering dimension and the $\ell_p$ loss (section 3), (2) Near-optimal approximate compression for linear functions and the $\ell_p$ loss (section 4.1) The results in sections 4.2 and 4.3 appear in the previous version