English

The Hamilton-Jacobi Theory of Deep Learning

Machine Learning 2026-05-29 v1 Artificial Intelligence Dynamical Systems Representation Theory Computational Physics

Abstract

In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator best fits the observations; at inference, the input is the spatial point at which that solution is evaluated and the initial condition is already encoded in the weights. The correspondence is exact for log-sum-exp layers and structural for broader architectures: residual networks, transformers, and recurrent architectures (RNNs, LSTMs, SSMs) each discretize the same class of Hamilton--Jacobi equations, with architecture-dependent Hamiltonian and viscosity. A single deformation parameter ε\varepsilon unifies all four perspectives (network, tropical algebra, viscous PDE, convex optimization) in a commutative diagram closed under Lipschitz conditions. Quantitative consequences include: the minimax optimal generalization rate O(n1/(d+2))O(n^{-1/(d+2)}) for fixed tt; adversarial robustness controlled by ε\varepsilon; backpropagation as the co-state equation of the Hamiltonian system for residual networks (Pontryagin Maximum Principle); scaling exponents consistent with data intrinsic dimension via PDE quadrature; and a closed-form O(N)O(N) influence function (softmax attribution weights πj\pi_j) whose entropy landscape undergoes fold bifurcations as ε\varepsilon increases, each merging attribution basins.

Keywords

Cite

@article{arxiv.2605.28983,
  title  = {The Hamilton-Jacobi Theory of Deep Learning},
  author = {Jose Marie Antonio Miñoza and Erika Fille T. Legara and Christopher P. Monterola},
  journal= {arXiv preprint arXiv:2605.28983},
  year   = {2026}
}