Sharp Stability Threshold and Certification for Designing Stable Residual Architectures
Abstract
We propose \emph{the sublinear-growth principle} for deep residual architectures -- a sharp stability threshold on the input-magnitude exponent of every residual block's velocity field: The threshold is established via two independent arguments. Classical ODE theory gives a global forward flow on at and exhibits divergent velocity fields at any . The optimal-control analysis, via the Hamilton-Jacobi-Bellman equation, sharpens this to a selection statement: the training optimum is bang-bang on the boundary of the admissible class, so the optimum at blows up while the optimum at is safe by construction. The exponent criterion is thereby a necessary and sufficient condition for stable training. It clarifies architectural placements that ensure the stability of training and inference, explaining, for instance, the stabilizing role of layer normalization. The sublinear-growth velocity fields form \emph{the right function space} on which forward dynamics, adjoint sensitivity, and architectural composition are all well-controlled. An arithmetic of input-magnitude exponents under the five operations that build residual blocks enables efficient certification of at the level of architectural primitives, in place of ad hoc trial and error in the search for stable neural architectural designs. A parameter-free modification reduces the supercritical Mamba block from to without layer normalization, demonstrating this point. Experiments on Mamba and PatchTST confirm that the variants train stably: the criterion is the input-magnitude exponent, not the presence of a normalization layer.
Cite
@article{arxiv.2607.14576,
title = {Sharp Stability Threshold and Certification for Designing Stable Residual Architectures},
author = {Hyemin Gu and Michael Tyrrell and Tuhin Sahai and Markos A. Katsoulakis},
journal= {arXiv preprint arXiv:2607.14576},
year = {2026}
}