Characterizing the Representational Capacity of Neural Processes

Robin Young
Proceedings of The 1st Symposium on Probabilistic Machine Learning, PMLR 327:256-289, 2026.

Abstract

What functions can Neural Processes represent? We analyze the representational capacity of popular NP architectures: Conditional Neural Processes (CNPs), Attentive Neural Processes (ANPs), Transformer Neural Processes (TNPs), and their latent variants. We prove these architectures form a strict hierarchy. CNP-representable functions are exactly those depending on finitely many expected features of the context distribution. ANPs strictly generalize CNPs via query-dependent reweighting, enabling kernel smoothers. ConvCNPs and ANPs are incomparable; each contains functions outside the other, separated by stationarity versus translation equivariance. TNPs with $L$ self-attention layers capture $L$-hop context interactions. For latent NPs, we show finite-dimensional latents provide coherent sampling but do not circumvent encoder limitations; matching GP posterior distributions requires latent dimension scaling with context size. These results provide a theoretical foundation for architecture selection based on task structure.

Cite this Paper


BibTeX
@InProceedings{pmlr-v327-young26a, title = {Characterizing the Representational Capacity of Neural Processes}, author = {Young, Robin}, booktitle = {Proceedings of The 1st Symposium on Probabilistic Machine Learning}, pages = {256--289}, year = {2026}, editor = {Swaroop, Siddharth and RĂ¼gamer, David and Kristiadi, Agustinus}, volume = {327}, series = {Proceedings of Machine Learning Research}, month = {05 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v327/main/assets/young26a/young26a.pdf}, url = {https://proceedings.mlr.press/v327/young26a.html}, abstract = { What functions can Neural Processes represent? We analyze the representational capacity of popular NP architectures: Conditional Neural Processes (CNPs), Attentive Neural Processes (ANPs), Transformer Neural Processes (TNPs), and their latent variants. We prove these architectures form a strict hierarchy. CNP-representable functions are exactly those depending on finitely many expected features of the context distribution. ANPs strictly generalize CNPs via query-dependent reweighting, enabling kernel smoothers. ConvCNPs and ANPs are incomparable; each contains functions outside the other, separated by stationarity versus translation equivariance. TNPs with $L$ self-attention layers capture $L$-hop context interactions. For latent NPs, we show finite-dimensional latents provide coherent sampling but do not circumvent encoder limitations; matching GP posterior distributions requires latent dimension scaling with context size. These results provide a theoretical foundation for architecture selection based on task structure. } }
Endnote
%0 Conference Paper %T Characterizing the Representational Capacity of Neural Processes %A Robin Young %B Proceedings of The 1st Symposium on Probabilistic Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Siddharth Swaroop %E David RĂ¼gamer %E Agustinus Kristiadi %F pmlr-v327-young26a %I PMLR %P 256--289 %U https://proceedings.mlr.press/v327/young26a.html %V 327 %X What functions can Neural Processes represent? We analyze the representational capacity of popular NP architectures: Conditional Neural Processes (CNPs), Attentive Neural Processes (ANPs), Transformer Neural Processes (TNPs), and their latent variants. We prove these architectures form a strict hierarchy. CNP-representable functions are exactly those depending on finitely many expected features of the context distribution. ANPs strictly generalize CNPs via query-dependent reweighting, enabling kernel smoothers. ConvCNPs and ANPs are incomparable; each contains functions outside the other, separated by stationarity versus translation equivariance. TNPs with $L$ self-attention layers capture $L$-hop context interactions. For latent NPs, we show finite-dimensional latents provide coherent sampling but do not circumvent encoder limitations; matching GP posterior distributions requires latent dimension scaling with context size. These results provide a theoretical foundation for architecture selection based on task structure.
APA
Young, R.. (2026). Characterizing the Representational Capacity of Neural Processes. Proceedings of The 1st Symposium on Probabilistic Machine Learning, in Proceedings of Machine Learning Research 327:256-289 Available from https://proceedings.mlr.press/v327/young26a.html.

Related Material