S2M-Net: Spectral-Spatial Mixing with Morphology-Aware Adaptive Loss for Medical Image Segmentation

Sanaullah Chowdhury, Lameya Sabrin
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:20589-20624, 2026.

Abstract

Medical image segmentation requires balancing global context with computational efficiency, where self-attention mechanisms suffer from quadratic $\mathcal{O}((HW)^2 C)$ complexity. We propose S2M-Net, a parameter-efficient architecture (4.7M parameters) that achieves computational savings through Spectral–Spatial Token Mixing (SSTM). SSTM achieves $\mathcal{O}(HWC^2)$ complexity through efficient combination of $\mathcal{O}(HWC \log(HW))$ frequency-domain processing and $\mathcal{O}(HWCd)$ bottlenecked spatial gating ($d{=}16$), exploiting spectral concentration where $>93%$ of energy is captured by $K{=}32$ low-frequency components ($\sim$0.8% of the spectrum at $352{\times}352$ resolution). This design avoids self-attention’s prohibitive $\mathcal{O}((HW)^2C)$ attention map computations while preserving global receptive fields. To handle geometric diversity, we introduce Morphology-Aware Adaptive Segmentation Loss (MASL), which automatically modulates five loss objectives based on per-sample morphological descriptors (tubularity, compactness, irregularity, and scale). Evaluation across 15 datasets spanning 8 modalities demonstrates competitive performance, obtaining the best performance on 14 of 15 datasets, with statistically significant improvements ($p < 0.0033$, Bonferroni-corrected) on 7 challenging tasks (complex morphology, class imbalance, and multi-class segmentation), and clinically meaningful gains ($0.5$–$1.6%$ Dice) on 8 mature benchmarks. Notably, S2M-Net achieves $83.43%$ Dice on EndoVis17 multiclass instrument segmentation ($+8.69%$ over TransUNet and $+9.14%$ over the best baseline UMamba at $74.29%$), while using $12.8{\times}$ fewer parameters (4.7M vs. 60M).

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chowdhury26b, title = {{S}2{M}-Net: Spectral-Spatial Mixing with Morphology-Aware Adaptive Loss for Medical Image Segmentation}, author = {Chowdhury, Sanaullah and Sabrin, Lameya}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {20589--20624}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chowdhury26b/chowdhury26b.pdf}, url = {https://proceedings.mlr.press/v306/chowdhury26b.html}, abstract = {Medical image segmentation requires balancing global context with computational efficiency, where self-attention mechanisms suffer from quadratic $\mathcal{O}((HW)^2 C)$ complexity. We propose S2M-Net, a parameter-efficient architecture (4.7M parameters) that achieves computational savings through Spectral–Spatial Token Mixing (SSTM). SSTM achieves $\mathcal{O}(HWC^2)$ complexity through efficient combination of $\mathcal{O}(HWC \log(HW))$ frequency-domain processing and $\mathcal{O}(HWCd)$ bottlenecked spatial gating ($d{=}16$), exploiting spectral concentration where $>93%$ of energy is captured by $K{=}32$ low-frequency components ($\sim$0.8% of the spectrum at $352{\times}352$ resolution). This design avoids self-attention’s prohibitive $\mathcal{O}((HW)^2C)$ attention map computations while preserving global receptive fields. To handle geometric diversity, we introduce Morphology-Aware Adaptive Segmentation Loss (MASL), which automatically modulates five loss objectives based on per-sample morphological descriptors (tubularity, compactness, irregularity, and scale). Evaluation across 15 datasets spanning 8 modalities demonstrates competitive performance, obtaining the best performance on 14 of 15 datasets, with statistically significant improvements ($p < 0.0033$, Bonferroni-corrected) on 7 challenging tasks (complex morphology, class imbalance, and multi-class segmentation), and clinically meaningful gains ($0.5$–$1.6%$ Dice) on 8 mature benchmarks. Notably, S2M-Net achieves $83.43%$ Dice on EndoVis17 multiclass instrument segmentation ($+8.69%$ over TransUNet and $+9.14%$ over the best baseline UMamba at $74.29%$), while using $12.8{\times}$ fewer parameters (4.7M vs. 60M).} }
Endnote
%0 Conference Paper %T S2M-Net: Spectral-Spatial Mixing with Morphology-Aware Adaptive Loss for Medical Image Segmentation %A Sanaullah Chowdhury %A Lameya Sabrin %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chowdhury26b %I PMLR %P 20589--20624 %U https://proceedings.mlr.press/v306/chowdhury26b.html %V 306 %X Medical image segmentation requires balancing global context with computational efficiency, where self-attention mechanisms suffer from quadratic $\mathcal{O}((HW)^2 C)$ complexity. We propose S2M-Net, a parameter-efficient architecture (4.7M parameters) that achieves computational savings through Spectral–Spatial Token Mixing (SSTM). SSTM achieves $\mathcal{O}(HWC^2)$ complexity through efficient combination of $\mathcal{O}(HWC \log(HW))$ frequency-domain processing and $\mathcal{O}(HWCd)$ bottlenecked spatial gating ($d{=}16$), exploiting spectral concentration where $>93%$ of energy is captured by $K{=}32$ low-frequency components ($\sim$0.8% of the spectrum at $352{\times}352$ resolution). This design avoids self-attention’s prohibitive $\mathcal{O}((HW)^2C)$ attention map computations while preserving global receptive fields. To handle geometric diversity, we introduce Morphology-Aware Adaptive Segmentation Loss (MASL), which automatically modulates five loss objectives based on per-sample morphological descriptors (tubularity, compactness, irregularity, and scale). Evaluation across 15 datasets spanning 8 modalities demonstrates competitive performance, obtaining the best performance on 14 of 15 datasets, with statistically significant improvements ($p < 0.0033$, Bonferroni-corrected) on 7 challenging tasks (complex morphology, class imbalance, and multi-class segmentation), and clinically meaningful gains ($0.5$–$1.6%$ Dice) on 8 mature benchmarks. Notably, S2M-Net achieves $83.43%$ Dice on EndoVis17 multiclass instrument segmentation ($+8.69%$ over TransUNet and $+9.14%$ over the best baseline UMamba at $74.29%$), while using $12.8{\times}$ fewer parameters (4.7M vs. 60M).
APA
Chowdhury, S. & Sabrin, L.. (2026). S2M-Net: Spectral-Spatial Mixing with Morphology-Aware Adaptive Loss for Medical Image Segmentation. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:20589-20624 Available from https://proceedings.mlr.press/v306/chowdhury26b.html.

Related Material