Three Costs of Amortizing Gaussian Process Inference with Neural Processes

Robin Young
Proceedings of the 2nd International Conference on Probabilistic Numerics, PMLR 341:227-244, 2026.

Abstract

Neural processes amortize Gaussian process inference, replacing the exact $O(n^3)$ posterior with a learned $O(n)$ map from context sets to predictive distributions. For a class of latent neural processes, we bound the Kullback–Leibler (KL) divergence between the GP and LNP predictives, decomposing it into three interpretable sources, namely label contamination as the neural process uses label values to estimate a quantity that is label-independent in the exact GP, an information bottleneck because the finite-dimensional representation cannot resolve the full context geometry, and amortization error from a single encoder network shared across all contexts. The bottleneck truncation term decays in the representation dimension $d$ as $O(e^{-cd^{2/d_x}})$ for squared-exponential kernels on $\mathbb{R}^{d_x}$ where $c > 0$ is a kernel-dependent constant and as $O(d^{-2\nu/d_x})$ for Matérn-$\nu$ kernels, directly linking architecture sizing to kernel smoothness and input dimension. The label contamination term is $O(1)$ in general, with only the observation-noise component decaying as $O(1/n)$, identifying a persistent cost of routing uncertainty estimation through a label-dependent representation. These results characterize the costs of amortization within the analyzed class and yield architectural recommendations to predict variance from context locations alone in the GP-amortization regime, and replace mean aggregation with second-order pooling to close the dominant amortization gap.

Cite this Paper


BibTeX
@InProceedings{pmlr-v341-young26a, title = {Three Costs of Amortizing {G}aussian Process Inference with Neural Processes}, author = {Young, Robin}, booktitle = {Proceedings of the 2nd International Conference on Probabilistic Numerics}, pages = {227--244}, year = {2026}, editor = {Karvonen, Toni and Bosch, Nathanael and Cockayne, Jon and Gessner, Alexandra and Hennig, Philipp and Kouw, Wouter}, volume = {341}, series = {Proceedings of Machine Learning Research}, month = {09--11 Sep}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v341/main/assets/young26a/young26a.pdf}, url = {https://proceedings.mlr.press/v341/young26a.html}, abstract = {Neural processes amortize Gaussian process inference, replacing the exact $O(n^3)$ posterior with a learned $O(n)$ map from context sets to predictive distributions. For a class of latent neural processes, we bound the Kullback–Leibler (KL) divergence between the GP and LNP predictives, decomposing it into three interpretable sources, namely label contamination as the neural process uses label values to estimate a quantity that is label-independent in the exact GP, an information bottleneck because the finite-dimensional representation cannot resolve the full context geometry, and amortization error from a single encoder network shared across all contexts. The bottleneck truncation term decays in the representation dimension $d$ as $O(e^{-cd^{2/d_x}})$ for squared-exponential kernels on $\mathbb{R}^{d_x}$ where $c > 0$ is a kernel-dependent constant and as $O(d^{-2\nu/d_x})$ for Matérn-$\nu$ kernels, directly linking architecture sizing to kernel smoothness and input dimension. The label contamination term is $O(1)$ in general, with only the observation-noise component decaying as $O(1/n)$, identifying a persistent cost of routing uncertainty estimation through a label-dependent representation. These results characterize the costs of amortization within the analyzed class and yield architectural recommendations to predict variance from context locations alone in the GP-amortization regime, and replace mean aggregation with second-order pooling to close the dominant amortization gap.} }
Endnote
%0 Conference Paper %T Three Costs of Amortizing Gaussian Process Inference with Neural Processes %A Robin Young %B Proceedings of the 2nd International Conference on Probabilistic Numerics %C Proceedings of Machine Learning Research %D 2026 %E Toni Karvonen %E Nathanael Bosch %E Jon Cockayne %E Alexandra Gessner %E Philipp Hennig %E Wouter Kouw %F pmlr-v341-young26a %I PMLR %P 227--244 %U https://proceedings.mlr.press/v341/young26a.html %V 341 %X Neural processes amortize Gaussian process inference, replacing the exact $O(n^3)$ posterior with a learned $O(n)$ map from context sets to predictive distributions. For a class of latent neural processes, we bound the Kullback–Leibler (KL) divergence between the GP and LNP predictives, decomposing it into three interpretable sources, namely label contamination as the neural process uses label values to estimate a quantity that is label-independent in the exact GP, an information bottleneck because the finite-dimensional representation cannot resolve the full context geometry, and amortization error from a single encoder network shared across all contexts. The bottleneck truncation term decays in the representation dimension $d$ as $O(e^{-cd^{2/d_x}})$ for squared-exponential kernels on $\mathbb{R}^{d_x}$ where $c > 0$ is a kernel-dependent constant and as $O(d^{-2\nu/d_x})$ for Matérn-$\nu$ kernels, directly linking architecture sizing to kernel smoothness and input dimension. The label contamination term is $O(1)$ in general, with only the observation-noise component decaying as $O(1/n)$, identifying a persistent cost of routing uncertainty estimation through a label-dependent representation. These results characterize the costs of amortization within the analyzed class and yield architectural recommendations to predict variance from context locations alone in the GP-amortization regime, and replace mean aggregation with second-order pooling to close the dominant amortization gap.
APA
Young, R.. (2026). Three Costs of Amortizing Gaussian Process Inference with Neural Processes. Proceedings of the 2nd International Conference on Probabilistic Numerics, in Proceedings of Machine Learning Research 341:227-244 Available from https://proceedings.mlr.press/v341/young26a.html.

Related Material