FSA-Bench: Benchmarking Federated Survival Models

Sultan Ahmed, Sanjay Purushotham
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:30-54, 2026.

Abstract

Clinical time-to-event data are routinely collected across multiple healthcare institutions but are rarely pooled due to privacy regulations such as GDPR and HIPAA. This fragmentation limits the development of robust and generalizable survival models for clinical decision-making. Federated Survival Analysis (FSA) enables collaborative modeling without sharing patient-level data, but the reliability, calibration, and robustness of survival models in federated settings remain insufficiently understood. We present FSA-Bench, a comprehensive benchmark for evaluating survival models under realistic federated and multi-institutional healthcare conditions. FSA-Bench includes classical statistical methods (Kaplan Meier, Weibull Accelerated Failure Time, and Cox Proportional Hazards), machine learning models, and modern deep learning approaches, including attention-based federated architectures such as FedPAttn. The benchmark evaluates these models across diverse clinical datasets under realistic challenges, including heterogeneous patient populations, varying censoring rates, and non-IID data distributed across multiple institutions. Across diverse clinical benchmark datasets, we evaluate model performance using clinically meaningful metrics, including discrimination (time-dependent C-index), calibration (Integrated Brier Score and negative log-likelihood), and ranking consistency using Spearman’s correlation, Kendall’s rank correlation, stability, and Jaccard similarity. An aggregated cross-dataset analysis further provides a comprehensive assessment of ranking robustness, revealing that Fed CoxPH offers the strongest overall ranking consistency, while FedPLSTM and FedPAttn achieve the highest stability and top-model agreement across heterogeneous federated settings. Collec- tively, FSA-Bench establishes a standardized and reproducible evaluation framework for benchmarking federated survival models and offers practical guidance for selecting reliable survival modeling approaches in privacy-preserving, multi-institutional healthcare environments.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-ahmed26a, title = {FSA-Bench: Benchmarking Federated Survival Models}, author = {Ahmed, Sultan and Purushotham, Sanjay}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {30--54}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/ahmed26a/ahmed26a.pdf}, url = {https://proceedings.mlr.press/v340/ahmed26a.html}, abstract = {Clinical time-to-event data are routinely collected across multiple healthcare institutions but are rarely pooled due to privacy regulations such as GDPR and HIPAA. This fragmentation limits the development of robust and generalizable survival models for clinical decision-making. Federated Survival Analysis (FSA) enables collaborative modeling without sharing patient-level data, but the reliability, calibration, and robustness of survival models in federated settings remain insufficiently understood. We present FSA-Bench, a comprehensive benchmark for evaluating survival models under realistic federated and multi-institutional healthcare conditions. FSA-Bench includes classical statistical methods (Kaplan Meier, Weibull Accelerated Failure Time, and Cox Proportional Hazards), machine learning models, and modern deep learning approaches, including attention-based federated architectures such as FedPAttn. The benchmark evaluates these models across diverse clinical datasets under realistic challenges, including heterogeneous patient populations, varying censoring rates, and non-IID data distributed across multiple institutions. Across diverse clinical benchmark datasets, we evaluate model performance using clinically meaningful metrics, including discrimination (time-dependent C-index), calibration (Integrated Brier Score and negative log-likelihood), and ranking consistency using Spearman’s correlation, Kendall’s rank correlation, stability, and Jaccard similarity. An aggregated cross-dataset analysis further provides a comprehensive assessment of ranking robustness, revealing that Fed CoxPH offers the strongest overall ranking consistency, while FedPLSTM and FedPAttn achieve the highest stability and top-model agreement across heterogeneous federated settings. Collec- tively, FSA-Bench establishes a standardized and reproducible evaluation framework for benchmarking federated survival models and offers practical guidance for selecting reliable survival modeling approaches in privacy-preserving, multi-institutional healthcare environments.} }
Endnote
%0 Conference Paper %T FSA-Bench: Benchmarking Federated Survival Models %A Sultan Ahmed %A Sanjay Purushotham %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-ahmed26a %I PMLR %P 30--54 %U https://proceedings.mlr.press/v340/ahmed26a.html %V 340 %X Clinical time-to-event data are routinely collected across multiple healthcare institutions but are rarely pooled due to privacy regulations such as GDPR and HIPAA. This fragmentation limits the development of robust and generalizable survival models for clinical decision-making. Federated Survival Analysis (FSA) enables collaborative modeling without sharing patient-level data, but the reliability, calibration, and robustness of survival models in federated settings remain insufficiently understood. We present FSA-Bench, a comprehensive benchmark for evaluating survival models under realistic federated and multi-institutional healthcare conditions. FSA-Bench includes classical statistical methods (Kaplan Meier, Weibull Accelerated Failure Time, and Cox Proportional Hazards), machine learning models, and modern deep learning approaches, including attention-based federated architectures such as FedPAttn. The benchmark evaluates these models across diverse clinical datasets under realistic challenges, including heterogeneous patient populations, varying censoring rates, and non-IID data distributed across multiple institutions. Across diverse clinical benchmark datasets, we evaluate model performance using clinically meaningful metrics, including discrimination (time-dependent C-index), calibration (Integrated Brier Score and negative log-likelihood), and ranking consistency using Spearman’s correlation, Kendall’s rank correlation, stability, and Jaccard similarity. An aggregated cross-dataset analysis further provides a comprehensive assessment of ranking robustness, revealing that Fed CoxPH offers the strongest overall ranking consistency, while FedPLSTM and FedPAttn achieve the highest stability and top-model agreement across heterogeneous federated settings. Collec- tively, FSA-Bench establishes a standardized and reproducible evaluation framework for benchmarking federated survival models and offers practical guidance for selecting reliable survival modeling approaches in privacy-preserving, multi-institutional healthcare environments.
APA
Ahmed, S. & Purushotham, S.. (2026). FSA-Bench: Benchmarking Federated Survival Models. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:30-54 Available from https://proceedings.mlr.press/v340/ahmed26a.html.

Related Material