English

Why and When Deep is Better than Shallow: Implementation-Agnostic State-Transition Model of Deep Learning

Machine Learning 2026-05-08 v4 Dynamical Systems Machine Learning

Abstract

Why and when does depth improve generalization? We study this question in an implementation-agnostic state-transition model, where a depth-kk predictor is a readout class HH composed with the word ball B(k,F)B(k,F) generated by hidden state transitions. Generalization bounds separate implementation error, approximation error, and statistical complexity, and upper bound the depth-dependent variance term by a Dudley entropy integral over B(k,F)B(k,F), with a conditional lower-bound diagnostic under readout separation. We identify geometric and semigroup mechanisms that keep this entropy contribution saturated or polynomial, and contrast them with separation mechanisms that recover the classical exponential-growth obstruction. Coupling these variance upper bounds with approximation rates gives typical depth trade-off patterns, clarifying that depth is statistically favorable when approximation improves rapidly while the transition semigroup remains geometrically tame.

Keywords

Cite

@article{arxiv.2505.15064,
  title  = {Why and When Deep is Better than Shallow: Implementation-Agnostic State-Transition Model of Deep Learning},
  author = {Sho Sonoda and Yuka Hashimoto and Isao Ishikawa and Masahiro Ikeda},
  journal= {arXiv preprint arXiv:2505.15064},
  year   = {2026}
}
R2 v1 2026-07-01T02:27:11.985Z