随机次梯度下降在一般可定义函数上收敛到极小点
最优化与控制
2022-02-14 v3 机器学习
摘要
Davis 和 Drusvyatskiy 先前已证明,一个一般的、半代数(更一般地在 o-极小结构中可定义)弱凸函数的每个 Clarke 临界点都位于一个活动流形上,并且要么是局部极小点,要么是活动严格鞍点。在本文第一部分,我们表明当弱凸性假设不成立时,会出现第三类点:尖锐排斥临界点。此外,我们证明了相应的活动流形满足我们在先前工作中引入的 Verdier 条件和角度条件。在本文第二部分,我们表明在扰动序列的一种密度型假设下,随机次梯度下降(SGD)以概率一避开尖锐排斥临界点。我们表明,通过在算法每次迭代中添加一个小的随机扰动(例如非退化高斯扰动),可以获得这样的密度型假设。这些结果结合我们先前关于避开活动严格鞍点的工作,表明在一般可定义(例如半代数)函数上的 SGD 收敛到局部极小点。
引用
@article{arxiv.2109.02455,
title = {Stochastic Subgradient Descent on a Generic Definable Function Converges to a Minimizer},
author = {Sholom Schechtman},
journal= {arXiv preprint arXiv:2109.02455},
year = {2022}
}
备注
This paper was withdrawn due to a mistake in the work of Bena\"im-Hofbauer-Sorin "Stochastic Approximations and Differential Inclusions". In the latter, the equivalence in Theorem 4.1 is not true and in particular the linearly interpolated process of the iterates is not an APT of the associated DI. This equivalence was at the heart of Propositions 7, 8 and Theorem 2 of the present paper