随机凸优化中梯度下降方法的信息论泛化界之局限
机器学习
2023-07-19 v3 机器学习
摘要
迄今为止,在随机凸优化设定下,尚无用于推断泛化误差的“信息论”框架被证明可确立梯度下降的极小极大速率。本工作中,我们考虑通过若干既有信息论框架确立此类速率的前景:输入输出互信息界、条件互信息界及其变体、PAC-Bayes 界以及近期的相关条件变体。我们证明这些界均无法确立极小极大速率。继而我们考虑研究梯度方法时常采用的一种策略,即对最终迭代施加高斯噪声扰动,产生带噪的“代理”算法。我们证明,通过对此类代理算法的分析亦无法确立极小极大速率。我们的结果表明,使用信息论技术分析梯度下降需要新的思路。
引用
@article{arxiv.2212.13556,
title = {Limitations of Information-Theoretic Generalization Bounds for Gradient Descent Methods in Stochastic Convex Optimization},
author = {Mahdi Haghifam and Borja Rodríguez-Gálvez and Ragnar Thobaben and Mikael Skoglund and Daniel M. Roy and Gintare Karolina Dziugaite},
journal= {arXiv preprint arXiv:2212.13556},
year = {2023}
}
备注
49 pages, 2 figures. This version corrects a mistake in the proof of Theorem 17. Proc. International Conference on Algorithmic Learning Theory (ALT), 2023