[edit]
FSA-Bench: Benchmarking Federated Survival Models
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:30-54, 2026.
Abstract
Clinical time-to-event data are routinely collected across multiple healthcare institutions but are rarely pooled due to privacy regulations such as GDPR and HIPAA. This fragmentation limits the development of robust and generalizable survival models for clinical decision-making. Federated Survival Analysis (FSA) enables collaborative modeling without sharing patient-level data, but the reliability, calibration, and robustness of survival models in federated settings remain insufficiently understood. We present FSA-Bench, a comprehensive benchmark for evaluating survival models under realistic federated and multi-institutional healthcare conditions. FSA-Bench includes classical statistical methods (Kaplan Meier, Weibull Accelerated Failure Time, and Cox Proportional Hazards), machine learning models, and modern deep learning approaches, including attention-based federated architectures such as FedPAttn. The benchmark evaluates these models across diverse clinical datasets under realistic challenges, including heterogeneous patient populations, varying censoring rates, and non-IID data distributed across multiple institutions. Across diverse clinical benchmark datasets, we evaluate model performance using clinically meaningful metrics, including discrimination (time-dependent C-index), calibration (Integrated Brier Score and negative log-likelihood), and ranking consistency using Spearman’s correlation, Kendall’s rank correlation, stability, and Jaccard similarity. An aggregated cross-dataset analysis further provides a comprehensive assessment of ranking robustness, revealing that Fed CoxPH offers the strongest overall ranking consistency, while FedPLSTM and FedPAttn achieve the highest stability and top-model agreement across heterogeneous federated settings. Collec- tively, FSA-Bench establishes a standardized and reproducible evaluation framework for benchmarking federated survival models and offers practical guidance for selecting reliable survival modeling approaches in privacy-preserving, multi-institutional healthcare environments.