English

Global Convergence of Over-parameterized Deep Equilibrium Models

Machine Learning 2023-03-30 v2 Machine Learning

Abstract

A deep equilibrium model (DEQ) is implicitly defined through an equilibrium point of an infinite-depth weight-tied model with an input-injection. Instead of infinite computations, it solves an equilibrium point directly with root-finding and computes gradients with implicit differentiation. The training dynamics of over-parameterized DEQs are investigated in this study. By supposing a condition on the initial equilibrium point, we show that the unique equilibrium point always exists during the training process, and the gradient descent is proved to converge to a globally optimal solution at a linear convergence rate for the quadratic loss function. In order to show that the required initial condition is satisfied via mild over-parameterization, we perform a fine-grained analysis on random DEQs. We propose a novel probabilistic framework to overcome the technical difficulty in the non-asymptotic analysis of infinite-depth weight-tied models.

Keywords

Cite

@article{arxiv.2205.13814,
  title  = {Global Convergence of Over-parameterized Deep Equilibrium Models},
  author = {Zenan Ling and Xingyu Xie and Qiuhao Wang and Zongpeng Zhang and Zhouchen Lin},
  journal= {arXiv preprint arXiv:2205.13814},
  year   = {2023}
}

Comments

Accepted by AISTATS 2023

R2 v1 2026-06-24T11:30:36.591Z