On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach

Edo Cohen-Karlik, Itamar Zimerman, Liane Galanti, Ido Andrew Atad, Amir Globerson, Lior Wolf
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:136-144, 2026.

Abstract

Recent advances in efficient sequence modeling have introduced selective state-space layers, a key component of the Mamba architecture, which have demonstrated remarkable success in a wide range of NLP and vision tasks. While Mamba’s empirical performance has matched or surpassed SoTA transformers on such diverse benchmarks, the theoretical foundations underlying its powerful representational capabilities remain less explored. In this work, we investigate the expressivity of selective state-space layers using multivariate polynomials, and prove that they surpass linear transformers in expressiveness. Consequently, our findings reveal that Mamba offers superior representational power over linear attention-based models for long-sequences, while not sacrificing their generalization. Our theoretical insights are validated by a comprehensive set of empirical experiments on various datasets.

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-cohen-karlik26a, title = { On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach }, author = {Cohen-Karlik, Edo and Zimerman, Itamar and Galanti, Liane and Atad, Ido Andrew and Globerson, Amir and Wolf, Lior}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {136--144}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/cohen-karlik26a/cohen-karlik26a.pdf}, url = {https://proceedings.mlr.press/v300/cohen-karlik26a.html}, abstract = { Recent advances in efficient sequence modeling have introduced selective state-space layers, a key component of the Mamba architecture, which have demonstrated remarkable success in a wide range of NLP and vision tasks. While Mamba’s empirical performance has matched or surpassed SoTA transformers on such diverse benchmarks, the theoretical foundations underlying its powerful representational capabilities remain less explored. In this work, we investigate the expressivity of selective state-space layers using multivariate polynomials, and prove that they surpass linear transformers in expressiveness. Consequently, our findings reveal that Mamba offers superior representational power over linear attention-based models for long-sequences, while not sacrificing their generalization. Our theoretical insights are validated by a comprehensive set of empirical experiments on various datasets. } }
Endnote
%0 Conference Paper %T On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach %A Edo Cohen-Karlik %A Itamar Zimerman %A Liane Galanti %A Ido Andrew Atad %A Amir Globerson %A Lior Wolf %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-cohen-karlik26a %I PMLR %P 136--144 %U https://proceedings.mlr.press/v300/cohen-karlik26a.html %V 300 %X Recent advances in efficient sequence modeling have introduced selective state-space layers, a key component of the Mamba architecture, which have demonstrated remarkable success in a wide range of NLP and vision tasks. While Mamba’s empirical performance has matched or surpassed SoTA transformers on such diverse benchmarks, the theoretical foundations underlying its powerful representational capabilities remain less explored. In this work, we investigate the expressivity of selective state-space layers using multivariate polynomials, and prove that they surpass linear transformers in expressiveness. Consequently, our findings reveal that Mamba offers superior representational power over linear attention-based models for long-sequences, while not sacrificing their generalization. Our theoretical insights are validated by a comprehensive set of empirical experiments on various datasets.
APA
Cohen-Karlik, E., Zimerman, I., Galanti, L., Atad, I.A., Globerson, A. & Wolf, L.. (2026). On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:136-144 Available from https://proceedings.mlr.press/v300/cohen-karlik26a.html.

Related Material