Representational Homomorphism Error Predicts Compositional Generalization In Language Models

Zhiyu An, Wan Du
Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations, PMLR 282:1-12, 2026.

Abstract

Compositional generalization—the ability to understand novel combinations of familiar components—remains a significant challenge for neural networks despite their success in many language tasks. Current evaluation methods focus on behavioral measures that reveal \emph{when} models fail to generalize compositionally, but provide limited insight into \emph{why} these failures occur at the representational level. We introduce \textit{Homomorphism Error} (HE), a structural metric that quantifies how well neural network representations preserve compositional operations by measuring deviations from approximate homomorphisms between expression spaces and their internal representations. Through controlled experiments on SCAN-style synthetic compositional tasks and small-scale Transformers, we demonstrate that HE serves as a strong predictor of out-of-distribution generalization performance, achieving $R^2 = 0.73$ correlation with OOD compositional generalization accuracy. Furthermore, our analysis reveals that model architecture has minimal impact on compositional structure, training data coverage exhibits threshold effects, but noise injection systematically degrades compositional representations in predictable ways. Importantly, we find that different aspects of compositionality—unary operations (modifiers) versus binary operations (sequence composition)—exhibit distinct sensitivities to distributional shifts, with modifier representations being particularly vulnerable to spurious correlations. These findings provide new mechanistic insights into compositional learning and establish homomorphism error as a valuable diagnostic tool for developing more robust neural architectures training methods. Code and data will be made publicaly available.

Cite this Paper


BibTeX
@InProceedings{pmlr-v282-an26a, title = {Representational Homomorphism Error Predicts Compositional Generalization In Language Models}, author = {An, Zhiyu and Du, Wan}, booktitle = {Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations}, pages = {1--12}, year = {2026}, editor = {Acosta, Francisco and Azeglio, Simone and Tolooshams, Bahareh and van de Geijn, Chase and Shewmake, Christian and Sanborn, Sophia and Miolane, Nina}, volume = {282}, series = {Proceedings of Machine Learning Research}, month = {14 Dec 2024--07 Dec 2025}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v282/main/assets/an26a/an26a.pdf}, url = {https://proceedings.mlr.press/v282/an26a.html}, abstract = {Compositional generalization—the ability to understand novel combinations of familiar components—remains a significant challenge for neural networks despite their success in many language tasks. Current evaluation methods focus on behavioral measures that reveal \emph{when} models fail to generalize compositionally, but provide limited insight into \emph{why} these failures occur at the representational level. We introduce \textit{Homomorphism Error} (HE), a structural metric that quantifies how well neural network representations preserve compositional operations by measuring deviations from approximate homomorphisms between expression spaces and their internal representations. Through controlled experiments on SCAN-style synthetic compositional tasks and small-scale Transformers, we demonstrate that HE serves as a strong predictor of out-of-distribution generalization performance, achieving $R^2 = 0.73$ correlation with OOD compositional generalization accuracy. Furthermore, our analysis reveals that model architecture has minimal impact on compositional structure, training data coverage exhibits threshold effects, but noise injection systematically degrades compositional representations in predictable ways. Importantly, we find that different aspects of compositionality—unary operations (modifiers) versus binary operations (sequence composition)—exhibit distinct sensitivities to distributional shifts, with modifier representations being particularly vulnerable to spurious correlations. These findings provide new mechanistic insights into compositional learning and establish homomorphism error as a valuable diagnostic tool for developing more robust neural architectures training methods. Code and data will be made publicaly available.} }
Endnote
%0 Conference Paper %T Representational Homomorphism Error Predicts Compositional Generalization In Language Models %A Zhiyu An %A Wan Du %B Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations %C Proceedings of Machine Learning Research %D 2026 %E Francisco Acosta %E Simone Azeglio %E Bahareh Tolooshams %E Chase van de Geijn %E Christian Shewmake %E Sophia Sanborn %E Nina Miolane %F pmlr-v282-an26a %I PMLR %P 1--12 %U https://proceedings.mlr.press/v282/an26a.html %V 282 %X Compositional generalization—the ability to understand novel combinations of familiar components—remains a significant challenge for neural networks despite their success in many language tasks. Current evaluation methods focus on behavioral measures that reveal \emph{when} models fail to generalize compositionally, but provide limited insight into \emph{why} these failures occur at the representational level. We introduce \textit{Homomorphism Error} (HE), a structural metric that quantifies how well neural network representations preserve compositional operations by measuring deviations from approximate homomorphisms between expression spaces and their internal representations. Through controlled experiments on SCAN-style synthetic compositional tasks and small-scale Transformers, we demonstrate that HE serves as a strong predictor of out-of-distribution generalization performance, achieving $R^2 = 0.73$ correlation with OOD compositional generalization accuracy. Furthermore, our analysis reveals that model architecture has minimal impact on compositional structure, training data coverage exhibits threshold effects, but noise injection systematically degrades compositional representations in predictable ways. Importantly, we find that different aspects of compositionality—unary operations (modifiers) versus binary operations (sequence composition)—exhibit distinct sensitivities to distributional shifts, with modifier representations being particularly vulnerable to spurious correlations. These findings provide new mechanistic insights into compositional learning and establish homomorphism error as a valuable diagnostic tool for developing more robust neural architectures training methods. Code and data will be made publicaly available.
APA
An, Z. & Du, W.. (2026). Representational Homomorphism Error Predicts Compositional Generalization In Language Models. Proceedings of the 4th (2025) and 3rd (2024) NeurIPS Workshops on Symmetry and Geometry in Neural Representations, in Proceedings of Machine Learning Research 282:1-12 Available from https://proceedings.mlr.press/v282/an26a.html.

Related Material