A Machine-Learned Comorbidity Index

Suleman Baloch, Kishlay Jha, Alberto Maria Segre, Philip M. Polgreen, Bijaya Adhikari
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:6041-6075, 2026.

Abstract

Traditional comorbidity scores (e.g., Charlson and Elixhauser) are widely used for risk adjustment and patient stratification, but they have two key limitations: (i) they are largely mortality-centric and do not align well with other clinical outcomes, and (ii) their linear, rule-based structure cannot capture nonlinear, outcome-specific risk relationships. We propose a Machine-Learned Comorbidity Index (MLCI) that maps diagnosis codes to a single scalar by maximizing the normalized Hilbert–Schmidt Independence Criterion (nHSIC) between the learned score and multiple clinical outcomes. MLCI captures nonlinear risk–outcome dependence and is supported by a theory that characterizes when a unified, informative admission-level ordering can be achieved across outcomes. Empirical results on multiple benchmark electronic health record (EHR) datasets show that MLCI outperforms strong baselines across multiple evaluation metrics.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-baloch26a, title = {A Machine-Learned Comorbidity Index}, author = {Baloch, Suleman and Jha, Kishlay and Segre, Alberto Maria and Polgreen, Philip M. and Adhikari, Bijaya}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {6041--6075}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/baloch26a/baloch26a.pdf}, url = {https://proceedings.mlr.press/v306/baloch26a.html}, abstract = {Traditional comorbidity scores (e.g., Charlson and Elixhauser) are widely used for risk adjustment and patient stratification, but they have two key limitations: (i) they are largely mortality-centric and do not align well with other clinical outcomes, and (ii) their linear, rule-based structure cannot capture nonlinear, outcome-specific risk relationships. We propose a Machine-Learned Comorbidity Index (MLCI) that maps diagnosis codes to a single scalar by maximizing the normalized Hilbert–Schmidt Independence Criterion (nHSIC) between the learned score and multiple clinical outcomes. MLCI captures nonlinear risk–outcome dependence and is supported by a theory that characterizes when a unified, informative admission-level ordering can be achieved across outcomes. Empirical results on multiple benchmark electronic health record (EHR) datasets show that MLCI outperforms strong baselines across multiple evaluation metrics.} }
Endnote
%0 Conference Paper %T A Machine-Learned Comorbidity Index %A Suleman Baloch %A Kishlay Jha %A Alberto Maria Segre %A Philip M. Polgreen %A Bijaya Adhikari %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-baloch26a %I PMLR %P 6041--6075 %U https://proceedings.mlr.press/v306/baloch26a.html %V 306 %X Traditional comorbidity scores (e.g., Charlson and Elixhauser) are widely used for risk adjustment and patient stratification, but they have two key limitations: (i) they are largely mortality-centric and do not align well with other clinical outcomes, and (ii) their linear, rule-based structure cannot capture nonlinear, outcome-specific risk relationships. We propose a Machine-Learned Comorbidity Index (MLCI) that maps diagnosis codes to a single scalar by maximizing the normalized Hilbert–Schmidt Independence Criterion (nHSIC) between the learned score and multiple clinical outcomes. MLCI captures nonlinear risk–outcome dependence and is supported by a theory that characterizes when a unified, informative admission-level ordering can be achieved across outcomes. Empirical results on multiple benchmark electronic health record (EHR) datasets show that MLCI outperforms strong baselines across multiple evaluation metrics.
APA
Baloch, S., Jha, K., Segre, A.M., Polgreen, P.M. & Adhikari, B.. (2026). A Machine-Learned Comorbidity Index. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:6041-6075 Available from https://proceedings.mlr.press/v306/baloch26a.html.

Related Material