Beyond Dice: Clinically Structured Bias Analysis of Multiple Sclerosis Lesion Segmentation Models

Abdul Basit, Namik Hassan, Muhammad Shafique
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:262-282, 2026.

Abstract

Automated multiple sclerosis (MS) lesion segmentation is commonly benchmarked with voxel-overlap metrics such as Dice, although clinical interpretation also depends on lesion size, anatomical location, and the preservation of subject-level dissemination evidence. We present *Beyond-Dice*, a reproducible MS-specific evaluation framework spanning voxel fidelity, lesion detection, anatomical location, subject-level evidence, case-type stress testing, and split/merge topology. We compare nnU-Net, SegResNet, and SwinUNETR on a separate 24-subject MSSEG-1 evaluation cohort containing 1,172 lesions. Subject-bootstrap 95% confidence intervals (CIs) show statistically tied Dice: 0.713 [0.648, 0.772] for nnU-Net, 0.713 [0.669, 0.759] for SegResNet, and 0.708 [0.656, 0.756] for SwinUNETR. Non-overlap endpoints nevertheless separate clinically relevant behavior: nnU-Net improves subject-mean lesion recall over SegResNet by 0.051 [0.016, 0.088], whereas SwinUNETR produces 10.67 additional false-positive lesions per scan relative to nnU-Net [5.88, 16.38]. Across all models, 238 lesions are missed by all three and 226 have model-dependent detection; lesion volume is the dominant detection predictor (odds ratio 3.86 per log-volume unit [3.45, 4.74]). A modality-matched MSLesSeg-to-MSSEG-1 experiment further shows similar matched-lesion Dice across locations (0.584–0.659) despite markedly lower infratentorial recall/F1 (0.339/0.412). These results operationalize how Dice-tied models can preserve different clinical evidence and support endpoint-specific, uncertainty-aware model selection rather than a single-score leaderboard.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-basit26a, title = {Beyond Dice: Clinically Structured Bias Analysis of Multiple Sclerosis Lesion Segmentation Models}, author = {Basit, Abdul and Hassan, Namik and Shafique, Muhammad}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {262--282}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/basit26a/basit26a.pdf}, url = {https://proceedings.mlr.press/v340/basit26a.html}, abstract = {Automated multiple sclerosis (MS) lesion segmentation is commonly benchmarked with voxel-overlap metrics such as Dice, although clinical interpretation also depends on lesion size, anatomical location, and the preservation of subject-level dissemination evidence. We present *Beyond-Dice*, a reproducible MS-specific evaluation framework spanning voxel fidelity, lesion detection, anatomical location, subject-level evidence, case-type stress testing, and split/merge topology. We compare nnU-Net, SegResNet, and SwinUNETR on a separate 24-subject MSSEG-1 evaluation cohort containing 1,172 lesions. Subject-bootstrap 95% confidence intervals (CIs) show statistically tied Dice: 0.713 [0.648, 0.772] for nnU-Net, 0.713 [0.669, 0.759] for SegResNet, and 0.708 [0.656, 0.756] for SwinUNETR. Non-overlap endpoints nevertheless separate clinically relevant behavior: nnU-Net improves subject-mean lesion recall over SegResNet by 0.051 [0.016, 0.088], whereas SwinUNETR produces 10.67 additional false-positive lesions per scan relative to nnU-Net [5.88, 16.38]. Across all models, 238 lesions are missed by all three and 226 have model-dependent detection; lesion volume is the dominant detection predictor (odds ratio 3.86 per log-volume unit [3.45, 4.74]). A modality-matched MSLesSeg-to-MSSEG-1 experiment further shows similar matched-lesion Dice across locations (0.584–0.659) despite markedly lower infratentorial recall/F1 (0.339/0.412). These results operationalize how Dice-tied models can preserve different clinical evidence and support endpoint-specific, uncertainty-aware model selection rather than a single-score leaderboard.} }
Endnote
%0 Conference Paper %T Beyond Dice: Clinically Structured Bias Analysis of Multiple Sclerosis Lesion Segmentation Models %A Abdul Basit %A Namik Hassan %A Muhammad Shafique %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-basit26a %I PMLR %P 262--282 %U https://proceedings.mlr.press/v340/basit26a.html %V 340 %X Automated multiple sclerosis (MS) lesion segmentation is commonly benchmarked with voxel-overlap metrics such as Dice, although clinical interpretation also depends on lesion size, anatomical location, and the preservation of subject-level dissemination evidence. We present *Beyond-Dice*, a reproducible MS-specific evaluation framework spanning voxel fidelity, lesion detection, anatomical location, subject-level evidence, case-type stress testing, and split/merge topology. We compare nnU-Net, SegResNet, and SwinUNETR on a separate 24-subject MSSEG-1 evaluation cohort containing 1,172 lesions. Subject-bootstrap 95% confidence intervals (CIs) show statistically tied Dice: 0.713 [0.648, 0.772] for nnU-Net, 0.713 [0.669, 0.759] for SegResNet, and 0.708 [0.656, 0.756] for SwinUNETR. Non-overlap endpoints nevertheless separate clinically relevant behavior: nnU-Net improves subject-mean lesion recall over SegResNet by 0.051 [0.016, 0.088], whereas SwinUNETR produces 10.67 additional false-positive lesions per scan relative to nnU-Net [5.88, 16.38]. Across all models, 238 lesions are missed by all three and 226 have model-dependent detection; lesion volume is the dominant detection predictor (odds ratio 3.86 per log-volume unit [3.45, 4.74]). A modality-matched MSLesSeg-to-MSSEG-1 experiment further shows similar matched-lesion Dice across locations (0.584–0.659) despite markedly lower infratentorial recall/F1 (0.339/0.412). These results operationalize how Dice-tied models can preserve different clinical evidence and support endpoint-specific, uncertainty-aware model selection rather than a single-score leaderboard.
APA
Basit, A., Hassan, N. & Shafique, M.. (2026). Beyond Dice: Clinically Structured Bias Analysis of Multiple Sclerosis Lesion Segmentation Models. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:262-282 Available from https://proceedings.mlr.press/v340/basit26a.html.

Related Material