TESLA: Taylor Expansion of Sinusoidal Learnable Activations

Daehwa Ko, JaeHyeon Kim, SeungHyun Ham, Jay Hoon Jung
Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, PMLR 300:5257-5265, 2026.

Abstract

The parity problem—deciding whether the number of ones in a binary vector is odd or even—remains challenging for standard neural networks due to linear inseparability and the need for global interactions. We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components. Theoretically, we show that constraining TESLA’s coefficients yields Lipschitz/Rademacher complexity bounds and shapes the training dynamics to emphasize higher-frequency structure. Empirically, on parity with input length $n=32$, TESLA attains strong generalization with 100K training samples ($\approx 0.002%$ of the $2^{32}$ input space) and remains robust under heavy corruption, retaining high accuracy with up to 30% label noise. We also compare against periodic and frequency-based baselines (SIREN, SNAKE, and Fourier feature embeddings) on parity and Forrelation. Beyond synthetic structure, TESLA delivers comparable performance on ImageNet-100, indicating that activation-level degree control transfers to more general vision workloads. Code: \url{https://github.com/KAU-QuantumAILab/TESLA}

Cite this Paper


BibTeX
@InProceedings{pmlr-v300-ko26a, title = { TESLA: Taylor Expansion of Sinusoidal Learnable Activations }, author = {Ko, Daehwa and Kim, JaeHyeon and Ham, SeungHyun and Jung, Jay Hoon}, booktitle = {Proceedings of The 29th International Conference on Artificial Intelligence and Statistics}, pages = {5257--5265}, year = {2026}, editor = {Khan, Emtiyaz and Li, Yingzhen and Solin, Arno and Ramdas, Aaditya}, volume = {300}, series = {Proceedings of Machine Learning Research}, month = {02--05 May}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v300/main/assets/ko26a/ko26a.pdf}, url = {https://proceedings.mlr.press/v300/ko26a.html}, abstract = { The parity problem—deciding whether the number of ones in a binary vector is odd or even—remains challenging for standard neural networks due to linear inseparability and the need for global interactions. We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components. Theoretically, we show that constraining TESLA’s coefficients yields Lipschitz/Rademacher complexity bounds and shapes the training dynamics to emphasize higher-frequency structure. Empirically, on parity with input length $n=32$, TESLA attains strong generalization with 100K training samples ($\approx 0.002%$ of the $2^{32}$ input space) and remains robust under heavy corruption, retaining high accuracy with up to 30% label noise. We also compare against periodic and frequency-based baselines (SIREN, SNAKE, and Fourier feature embeddings) on parity and Forrelation. Beyond synthetic structure, TESLA delivers comparable performance on ImageNet-100, indicating that activation-level degree control transfers to more general vision workloads. Code: \url{https://github.com/KAU-QuantumAILab/TESLA} } }
Endnote
%0 Conference Paper %T TESLA: Taylor Expansion of Sinusoidal Learnable Activations %A Daehwa Ko %A JaeHyeon Kim %A SeungHyun Ham %A Jay Hoon Jung %B Proceedings of The 29th International Conference on Artificial Intelligence and Statistics %C Proceedings of Machine Learning Research %D 2026 %E Emtiyaz Khan %E Yingzhen Li %E Arno Solin %E Aaditya Ramdas %F pmlr-v300-ko26a %I PMLR %P 5257--5265 %U https://proceedings.mlr.press/v300/ko26a.html %V 300 %X The parity problem—deciding whether the number of ones in a binary vector is odd or even—remains challenging for standard neural networks due to linear inseparability and the need for global interactions. We propose TESLA, an activation defined as a learnable combination of sine and cosine terms, enabling explicit control over polynomial degree and selective amplification of high-order components. Theoretically, we show that constraining TESLA’s coefficients yields Lipschitz/Rademacher complexity bounds and shapes the training dynamics to emphasize higher-frequency structure. Empirically, on parity with input length $n=32$, TESLA attains strong generalization with 100K training samples ($\approx 0.002%$ of the $2^{32}$ input space) and remains robust under heavy corruption, retaining high accuracy with up to 30% label noise. We also compare against periodic and frequency-based baselines (SIREN, SNAKE, and Fourier feature embeddings) on parity and Forrelation. Beyond synthetic structure, TESLA delivers comparable performance on ImageNet-100, indicating that activation-level degree control transfers to more general vision workloads. Code: \url{https://github.com/KAU-QuantumAILab/TESLA}
APA
Ko, D., Kim, J., Ham, S. & Jung, J.H.. (2026). TESLA: Taylor Expansion of Sinusoidal Learnable Activations . Proceedings of The 29th International Conference on Artificial Intelligence and Statistics, in Proceedings of Machine Learning Research 300:5257-5265 Available from https://proceedings.mlr.press/v300/ko26a.html.

Related Material