Understanding Task Representations in Neural Networks via Bayesian Ablation

Andrew Joohun Nam, Declan Iain Campbell, Thomas L. Griffiths, Jonathan D. Cohen, Sarah-Jane Leslie
Proceedings of the Fifth Conference on Causal Learning and Reasoning, PMLR 323:196-221, 2026.

Abstract

Neural networks are powerful tools for cognitive modeling due to their flexibility and emergent properties. However, interpreting their learned representations remains challenging due to their sub-symbolic semantics. We introduce a novel probabilistic framework for interpreting latent task representations in neural networks. Inspired by Bayesian inference, our approach defines a distribution over representational units to infer their causal contributions to task performance. Using ideas from information theory, we propose a suite of tools and metrics to illuminate key model properties, including representational distributedness, manifold complexity, and polysemanticity.

Cite this Paper


BibTeX
@InProceedings{pmlr-v323-nam26a, title = {Understanding Task Representations in Neural Networks via Bayesian Ablation}, author = {Nam, Andrew Joohun and Campbell, Declan Iain and Griffiths, Thomas L. and Cohen, Jonathan D. and Leslie, Sarah-Jane}, booktitle = {Proceedings of the Fifth Conference on Causal Learning and Reasoning}, pages = {196--221}, year = {2026}, editor = {Mazaheri, Bijan and Hanson, Niels Richard}, volume = {323}, series = {Proceedings of Machine Learning Research}, month = {06--08 Apr}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v323/main/assets/nam26a/nam26a.pdf}, url = {https://proceedings.mlr.press/v323/nam26a.html}, abstract = {Neural networks are powerful tools for cognitive modeling due to their flexibility and emergent properties. However, interpreting their learned representations remains challenging due to their sub-symbolic semantics. We introduce a novel probabilistic framework for interpreting latent task representations in neural networks. Inspired by Bayesian inference, our approach defines a distribution over representational units to infer their causal contributions to task performance. Using ideas from information theory, we propose a suite of tools and metrics to illuminate key model properties, including representational distributedness, manifold complexity, and polysemanticity.} }
Endnote
%0 Conference Paper %T Understanding Task Representations in Neural Networks via Bayesian Ablation %A Andrew Joohun Nam %A Declan Iain Campbell %A Thomas L. Griffiths %A Jonathan D. Cohen %A Sarah-Jane Leslie %B Proceedings of the Fifth Conference on Causal Learning and Reasoning %C Proceedings of Machine Learning Research %D 2026 %E Bijan Mazaheri %E Niels Richard Hanson %F pmlr-v323-nam26a %I PMLR %P 196--221 %U https://proceedings.mlr.press/v323/nam26a.html %V 323 %X Neural networks are powerful tools for cognitive modeling due to their flexibility and emergent properties. However, interpreting their learned representations remains challenging due to their sub-symbolic semantics. We introduce a novel probabilistic framework for interpreting latent task representations in neural networks. Inspired by Bayesian inference, our approach defines a distribution over representational units to infer their causal contributions to task performance. Using ideas from information theory, we propose a suite of tools and metrics to illuminate key model properties, including representational distributedness, manifold complexity, and polysemanticity.
APA
Nam, A.J., Campbell, D.I., Griffiths, T.L., Cohen, J.D. & Leslie, S.. (2026). Understanding Task Representations in Neural Networks via Bayesian Ablation. Proceedings of the Fifth Conference on Causal Learning and Reasoning, in Proceedings of Machine Learning Research 323:196-221 Available from https://proceedings.mlr.press/v323/nam26a.html.

Related Material