English

A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks

Machine Learning 2026-04-02 v3 Machine Learning

Abstract

Performing gradient descent in a wide neural network is equivalent to computing the posterior mean of a Gaussian Process with the Neural Tangent Kernel (NTK-GP), for a specific prior mean and with zero observation noise. However, existing formulations have two limitations: (i) the NTK-GP assumes noiseless targets, leading to misspecification on noisy data; (ii) the equivalence does not extend to arbitrary prior means, which are essential for well-specified models. To address (i), we introduce a regularizer into the training objective, showing its correspondence to incorporating observation noise in the NTK-GP. To address (ii), we propose a \textit{shifted network} that enables arbitrary prior means and allows obtaining the posterior mean with gradient descent on a single network, without ensembling or kernel inversion. We validate our results with experiments across datasets and architectures, showing that this approach removes key obstacles to the practical use of NTK-GP equivalence in applied Gaussian process modeling.

Keywords

Cite

@article{arxiv.2502.01556,
  title  = {A Gaussian Process View on Observation Noise and Initialization in Wide Neural Networks},
  author = {Sergio Calvo-Ordoñez and Jonathan Plenk and Richard Bergna and Alvaro Cartea and Jose Miguel Hernandez-Lobato and Konstantina Palla and Kamil Ciosek},
  journal= {arXiv preprint arXiv:2502.01556},
  year   = {2026}
}

Comments

AISTATS 2026, Camera-ready version