Some Theoretical Results on Layerwise Effective Dimension Oscillations in Finite Width ReLU Networks
Abstract
We analyze the layerwise effective dimension (rank of the feature matrix) in fully-connected ReLU networks of finite width. Specifically, for a fixed batch of inputs and random Gaussian weights, we derive closed-form expressions for the expected rank of the $m\times n$ hidden activation matrices. Our main result shows that so that the rank deficit decays geometrically with ratio . We also prove a sub-Gaussian concentration bound, and identify the "revival" depths at which the expected rank attains local maxima. In particular, these peaks occur at depths with height . We further show that this oscillatory rank behavior is a finite-width phenomenon: under orthogonal weight initialization or strong negative-slope leaky-ReLU, the rank remains (nearly) full. These results provide a precise characterization of how random ReLU layers alternately collapse and partially revive the subspace of input variations, adding nuance to prior work on expressivity of deep networks.
Keywords
Cite
@article{arxiv.2507.07675,
title = {Some Theoretical Results on Layerwise Effective Dimension Oscillations in Finite Width ReLU Networks},
author = {Darshan Makwana},
journal= {arXiv preprint arXiv:2507.07675},
year = {2025}
}
Comments
Incomplete citations