English

Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis

Machine Learning 2024-10-11 v2 Optimization and Control Machine Learning

Abstract

We analyze recurrent neural networks with diagonal hidden-to-hidden weight matrices, trained with gradient descent in the supervised learning setting, and prove that gradient descent can achieve optimality \emph{without} massive overparameterization. Our in-depth nonasymptotic analysis (i) provides improved bounds on the network size mm in terms of the sequence length TT, sample size nn and ambient dimension dd, and (ii) identifies the significant impact of long-term dependencies in the dynamical system on the convergence and network width bounds characterized by a cutoff point that depends on the Lipschitz continuity of the activation function. Remarkably, this analysis reveals that an appropriately-initialized recurrent neural network trained with nn samples can achieve optimality with a network size mm that scales only logarithmically with nn. This sharply contrasts with the prior works that require high-order polynomial dependency of mm on nn to establish strong regularity conditions. Our results are based on an explicit characterization of the class of dynamical systems that can be approximated and learned by recurrent neural networks via norm-constrained transportation mappings, and establishing local smoothness properties of the hidden state with respect to the learnable parameters.

Keywords

Cite

@article{arxiv.2402.12241,
  title  = {Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis},
  author = {Semih Cayci and Atilla Eryilmaz},
  journal= {arXiv preprint arXiv:2402.12241},
  year   = {2024}
}