基于压缩方案的 Gaussian 混合模型鲁棒学习的近最优样本复杂度界
机器学习
2020-07-23 v5 统计理论
统计理论
摘要
我们证明了 个样本对于学习 中 个 Gaussian 的混合是必要且充分的,其误差在总变差距离上为 。这改进了该问题已知的上界和下界。对于轴对齐 Gaussian 的混合,我们证明了 个样本即可满足要求,与已知的下界相匹配。此外,这些结果在不可知学习/鲁棒估计设定下同样成立,其中目标分布仅近似为 Gaussian 混合。上界是使用一种基于“压缩”概念的分布学习新技术来证明的。任何允许此类压缩方案的分布类也可以用少量样本进行学习。此外,如果一类分布具有这样的压缩方案,那么这些分布的乘积类和混合类也具有压缩方案。我们主要结果的核心是证明 中的 Gaussian 类允许小规模压缩方案。
引用
@article{arxiv.1710.05209,
title = {Near-optimal Sample Complexity Bounds for Robust Learning of Gaussians Mixtures via Compression Schemes},
author = {Hassan Ashtiani and Shai Ben-David and Nick Harvey and Christopher Liaw and Abbas Mehrabian and Yaniv Plan},
journal= {arXiv preprint arXiv:1710.05209},
year = {2020}
}
备注
To appear in Journal of the ACM. 46 pages. An extended abstract appeared in NeurIPS 2018. This version contains all the proofs, generalizes the results to agnostic learning, and improves the bounds by logarithmic factors