English

Convergence Rates of Training Deep Neural Networks via Alternating Minimization Methods

Machine Learning 2023-04-05 v2 Optimization and Control

Abstract

Training deep neural networks (DNNs) is an important and challenging optimization problem in machine learning due to its non-convexity and non-separable structure. The alternating minimization (AM) approaches split the composition structure of DNNs and have drawn great interest in the deep learning and optimization communities. In this paper, we propose a unified framework for analyzing the convergence rate of AM-type network training methods. Our analysis is based on the non-monotone jj-step sufficient decrease conditions and the Kurdyka-Lojasiewicz (KL) property, which relaxes the requirement of designing descent algorithms. We show the detailed local convergence rate if the KL exponent θ\theta varies in [0,1)[0,1). Moreover, the local R-linear convergence is discussed under a stronger jj-step sufficient decrease condition.

Keywords

Cite

@article{arxiv.2208.14318,
  title  = {Convergence Rates of Training Deep Neural Networks via Alternating Minimization Methods},
  author = {Jintao Xu and Chenglong Bao and Wenxun Xing},
  journal= {arXiv preprint arXiv:2208.14318},
  year   = {2023}
}