Three Costs of Amortizing Gaussian Process Inference with Neural Processes
Abstract
Neural processes amortize Gaussian process inference, replacing the exact posterior with a learned map from context sets to predictive distributions. For a class of latent neural processes, we bound the Kullback--Leibler (KL) divergence between the GP and LNP predictives, decomposing it into three interpretable sources, namely label contamination as the neural process uses label values to estimate a quantity that is label-independent in the exact GP, an information bottleneck because the finite-dimensional representation cannot resolve the full context geometry, and amortization error from a single encoder network shared across all contexts. The bottleneck truncation term decays in the representation dimension as for squared-exponential kernels on where is a kernel-dependent constant and as for Mat\'ern- kernels, directly linking architecture sizing to kernel smoothness and input dimension. The label contamination term is in general, with only the observation-noise component decaying as , identifying a persistent cost of routing uncertainty estimation through a label-dependent representation. These results characterize the costs of amortization within the analyzed class and yield architectural recommendations to predict variance from context locations alone in the GP-amortization regime, and replace mean aggregation with second-order pooling to close the dominant amortization gap.
Cite
@article{arxiv.2605.21798,
title = {Three Costs of Amortizing Gaussian Process Inference with Neural Processes},
author = {Robin Young},
journal= {arXiv preprint arXiv:2605.21798},
year = {2026}
}
Comments
To appear at ProbNum 2026