[edit]
Beyond Dice: Clinically Structured Bias Analysis of Multiple Sclerosis Lesion Segmentation Models
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:262-282, 2026.
Abstract
Automated multiple sclerosis (MS) lesion segmentation is commonly benchmarked with voxel-overlap metrics such as Dice, although clinical interpretation also depends on lesion size, anatomical location, and the preservation of subject-level dissemination evidence. We present *Beyond-Dice*, a reproducible MS-specific evaluation framework spanning voxel fidelity, lesion detection, anatomical location, subject-level evidence, case-type stress testing, and split/merge topology. We compare nnU-Net, SegResNet, and SwinUNETR on a separate 24-subject MSSEG-1 evaluation cohort containing 1,172 lesions. Subject-bootstrap 95% confidence intervals (CIs) show statistically tied Dice: 0.713 [0.648, 0.772] for nnU-Net, 0.713 [0.669, 0.759] for SegResNet, and 0.708 [0.656, 0.756] for SwinUNETR. Non-overlap endpoints nevertheless separate clinically relevant behavior: nnU-Net improves subject-mean lesion recall over SegResNet by 0.051 [0.016, 0.088], whereas SwinUNETR produces 10.67 additional false-positive lesions per scan relative to nnU-Net [5.88, 16.38]. Across all models, 238 lesions are missed by all three and 226 have model-dependent detection; lesion volume is the dominant detection predictor (odds ratio 3.86 per log-volume unit [3.45, 4.74]). A modality-matched MSLesSeg-to-MSSEG-1 experiment further shows similar matched-lesion Dice across locations (0.584–0.659) despite markedly lower infratentorial recall/F1 (0.339/0.412). These results operationalize how Dice-tied models can preserve different clinical evidence and support endpoint-specific, uncertainty-aware model selection rather than a single-score leaderboard.