When Can We Trust Survival Model Evaluation ?

Ghanem Bahrini, Sebastien Razakarivony, Jean-François Dupuy, Valerie Gares, Morgane Barbet-Massin
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:5219-5244, 2026.

Abstract

Evaluating survival models under censoring is inherently challenging, yet standard evaluation practices are often applied without explicitly assessing how censoring distorts metric reliability. Performing a large experimental study, we analyze and quantify how survival evaluation metrics are affected in fundamentally different ways by the censoring rate and the censoring mechanism. Using a controlled semi-synthetic framework, we vary both the censoring mechanism (administrative, independent, covariate-dependent) and the censoring rate, and compare standard evaluations based on censored data with oracle evaluations using fully observed event times. This controlled setting enables us to quantify distortions along two complementary axes: numerical bias and preservation of model ranking. Across datasets and metric families, we find that censoring induces systematic, mechanism-dependent distortions. Moderate numerical bias, if not properly addressed, can lead to unreliable model comparison as censoring increases. These findings reveal fundamental limitations of common benchmarking practices and call for more careful interpretation of survival evaluation under realistic censoring.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-bahrini26a, title = {When Can We Trust Survival Model Evaluation ?}, author = {Bahrini, Ghanem and Razakarivony, Sebastien and Dupuy, Jean-Fran\c{c}ois and Gares, Valerie and Barbet-Massin, Morgane}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {5219--5244}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/bahrini26a/bahrini26a.pdf}, url = {https://proceedings.mlr.press/v306/bahrini26a.html}, abstract = {Evaluating survival models under censoring is inherently challenging, yet standard evaluation practices are often applied without explicitly assessing how censoring distorts metric reliability. Performing a large experimental study, we analyze and quantify how survival evaluation metrics are affected in fundamentally different ways by the censoring rate and the censoring mechanism. Using a controlled semi-synthetic framework, we vary both the censoring mechanism (administrative, independent, covariate-dependent) and the censoring rate, and compare standard evaluations based on censored data with oracle evaluations using fully observed event times. This controlled setting enables us to quantify distortions along two complementary axes: numerical bias and preservation of model ranking. Across datasets and metric families, we find that censoring induces systematic, mechanism-dependent distortions. Moderate numerical bias, if not properly addressed, can lead to unreliable model comparison as censoring increases. These findings reveal fundamental limitations of common benchmarking practices and call for more careful interpretation of survival evaluation under realistic censoring.} }
Endnote
%0 Conference Paper %T When Can We Trust Survival Model Evaluation ? %A Ghanem Bahrini %A Sebastien Razakarivony %A Jean-François Dupuy %A Valerie Gares %A Morgane Barbet-Massin %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-bahrini26a %I PMLR %P 5219--5244 %U https://proceedings.mlr.press/v306/bahrini26a.html %V 306 %X Evaluating survival models under censoring is inherently challenging, yet standard evaluation practices are often applied without explicitly assessing how censoring distorts metric reliability. Performing a large experimental study, we analyze and quantify how survival evaluation metrics are affected in fundamentally different ways by the censoring rate and the censoring mechanism. Using a controlled semi-synthetic framework, we vary both the censoring mechanism (administrative, independent, covariate-dependent) and the censoring rate, and compare standard evaluations based on censored data with oracle evaluations using fully observed event times. This controlled setting enables us to quantify distortions along two complementary axes: numerical bias and preservation of model ranking. Across datasets and metric families, we find that censoring induces systematic, mechanism-dependent distortions. Moderate numerical bias, if not properly addressed, can lead to unreliable model comparison as censoring increases. These findings reveal fundamental limitations of common benchmarking practices and call for more careful interpretation of survival evaluation under realistic censoring.
APA
Bahrini, G., Razakarivony, S., Dupuy, J., Gares, V. & Barbet-Massin, M.. (2026). When Can We Trust Survival Model Evaluation ?. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:5219-5244 Available from https://proceedings.mlr.press/v306/bahrini26a.html.

Related Material