English

How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs

Machine Learning 2026-02-06 v1 Information Theory math.IT

Abstract

The intermediate layers of deep networks can be characterised as a Gaussian process, in particular the Edge-of-Chaos (EoC) initialisation strategy prescribes the limiting covariance matrix of the Gaussian process. Here we show that the under-utilised chosen variance of the Gaussian process is important in the training of deep networks with sparsity inducing activation, such as a shifted and clipped ReLU, CReLUτ,m(x)=min(max(xτ,0),m)\text{CReLU}_{\tau,m}(x)=\min(\max(x-\tau,0),m). Specifically, initialisations leading to larger fixed Gaussian process variances, allow for improved expressivity with activation sparsity as large as 90% in DNNs and CNNs, and generally improve the stability of the training process. Enabling full, or near full, accuracy at such high levels of sparsity in the hidden layers suggests a promising mechanism to reduce the energy consumption of machine learning models involving fully connected layers.

Keywords

Cite

@article{arxiv.2602.05779,
  title  = {How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs},
  author = {Emily Dent and Jared Tanner},
  journal= {arXiv preprint arXiv:2602.05779},
  year   = {2026}
}