English

Injectivity of ReLU networks: perspectives from statistical physics

Disordered Systems and Neural Networks 2024-12-13 v2 Machine Learning Probability Machine Learning

Abstract

When can the input of a ReLU neural network be inferred from its output? In other words, when is the network injective? We consider a single layer, xReLU(Wx)x \mapsto \mathrm{ReLU}(Wx), with a random Gaussian m×nm \times n matrix WW, in a high-dimensional setting where n,mn, m \to \infty. Recent work connects this problem to spherical integral geometry giving rise to a conjectured sharp injectivity threshold for α=mn\alpha = \frac{m}{n} by studying the expected Euler characteristic of a certain random set. We adopt a different perspective and show that injectivity is equivalent to a property of the ground state of the spherical perceptron, an important spin glass model in statistical physics. By leveraging the (non-rigorous) replica symmetry-breaking theory, we derive analytical equations for the threshold whose solution is at odds with that from the Euler characteristic. Furthermore, we use Gordon's min--max theorem to prove that a replica-symmetric upper bound refutes the Euler characteristic prediction. Along the way we aim to give a tutorial-style introduction to key ideas from statistical physics in an effort to make the exposition accessible to a broad audience. Our analysis establishes a connection between spin glasses and integral geometry but leaves open the problem of explaining the discrepancies.

Cite

@article{arxiv.2302.14112,
  title  = {Injectivity of ReLU networks: perspectives from statistical physics},
  author = {Antoine Maillard and Afonso S. Bandeira and David Belius and Ivan Dokmanić and Shuta Nakajima},
  journal= {arXiv preprint arXiv:2302.14112},
  year   = {2024}
}

Comments

62 pages ; Changes to match the published version (v2), in particular Appendix A.7 was added, and Appendix G was re-worked as an alternative proof of Theorem 1.8

R2 v1 2026-06-28T08:51:03.984Z