Caracal: Causal Architecture via Spectral Mixing

Bingzheng Gan, Tianyi Zhang, Li Yusu, Jing Huang, Wei Shi, Yangkai Ding, Tao Yu
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:32876-32895, 2026.

Abstract

The scalability of Large Language Models to long sequences is hindered by the quadratic cost of self-attention and the limitations of positional encodings. To address these, we introduce Caracal, a novel architecture that replaces self-attention with a parameter-efficient, $\mathcal{O}(L \log L)$ Multi-Head Fourier (MHF) module. Our contributions are threefold: (1) We leverage the Fast Fourier Transform (FFT) for sequence mixing, inherently addressing both bottlenecks mentioned above. (2) We apply a frequency-domain causal masking technique that enforces autoregressive capabilities via asymmetric padding and truncation, overcoming a critical barrier for Fourier-based generative models. (3) Unlike efficient models relying on hardware-specific implementations (e.g., Mamba), Caracal uses standard library operators. This ensures robust portability, eliminating common deployment barriers. Evaluations demonstrate that Caracal performs competitively with Transformer and SSM baselines, offering a scalable and simple pathway for efficient long-sequence modeling. Code is available in the supplementary materials.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-gan26c, title = {Caracal: Causal Architecture via Spectral Mixing}, author = {Gan, Bingzheng and Zhang, Tianyi and Yusu, Li and Huang, Jing and Shi, Wei and Ding, Yangkai and Yu, Tao}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {32876--32895}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/gan26c/gan26c.pdf}, url = {https://proceedings.mlr.press/v306/gan26c.html}, abstract = {The scalability of Large Language Models to long sequences is hindered by the quadratic cost of self-attention and the limitations of positional encodings. To address these, we introduce Caracal, a novel architecture that replaces self-attention with a parameter-efficient, $\mathcal{O}(L \log L)$ Multi-Head Fourier (MHF) module. Our contributions are threefold: (1) We leverage the Fast Fourier Transform (FFT) for sequence mixing, inherently addressing both bottlenecks mentioned above. (2) We apply a frequency-domain causal masking technique that enforces autoregressive capabilities via asymmetric padding and truncation, overcoming a critical barrier for Fourier-based generative models. (3) Unlike efficient models relying on hardware-specific implementations (e.g., Mamba), Caracal uses standard library operators. This ensures robust portability, eliminating common deployment barriers. Evaluations demonstrate that Caracal performs competitively with Transformer and SSM baselines, offering a scalable and simple pathway for efficient long-sequence modeling. Code is available in the supplementary materials.} }
Endnote
%0 Conference Paper %T Caracal: Causal Architecture via Spectral Mixing %A Bingzheng Gan %A Tianyi Zhang %A Li Yusu %A Jing Huang %A Wei Shi %A Yangkai Ding %A Tao Yu %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-gan26c %I PMLR %P 32876--32895 %U https://proceedings.mlr.press/v306/gan26c.html %V 306 %X The scalability of Large Language Models to long sequences is hindered by the quadratic cost of self-attention and the limitations of positional encodings. To address these, we introduce Caracal, a novel architecture that replaces self-attention with a parameter-efficient, $\mathcal{O}(L \log L)$ Multi-Head Fourier (MHF) module. Our contributions are threefold: (1) We leverage the Fast Fourier Transform (FFT) for sequence mixing, inherently addressing both bottlenecks mentioned above. (2) We apply a frequency-domain causal masking technique that enforces autoregressive capabilities via asymmetric padding and truncation, overcoming a critical barrier for Fourier-based generative models. (3) Unlike efficient models relying on hardware-specific implementations (e.g., Mamba), Caracal uses standard library operators. This ensures robust portability, eliminating common deployment barriers. Evaluations demonstrate that Caracal performs competitively with Transformer and SSM baselines, offering a scalable and simple pathway for efficient long-sequence modeling. Code is available in the supplementary materials.
APA
Gan, B., Zhang, T., Yusu, L., Huang, J., Shi, W., Ding, Y. & Yu, T.. (2026). Caracal: Causal Architecture via Spectral Mixing. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:32876-32895 Available from https://proceedings.mlr.press/v306/gan26c.html.

Related Material