Censoring-Aware Reinforcement Learning to Optimize Early Risk Alerts from Longitudinal Clinical Data

Qin Weng, Benjamin Goldstein, Matthew M. Engelhard
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:2151-2174, 2026.

Abstract

Early recognition of chronic conditions is critical to ensure patients receive timely interventions and support. Passive surveillance of routine electronic health records (EHRs) provides information about longitudinal health trajectories that can support prompt recognition and inform associated early actions. However, relevant information is acquired at a different rate for each patient, and there is an inherent trade-off between the earliness versus the specificity of diagnosis and related actions. Therefore, determining when to alert providers about a likely chronic condition requires us to weigh the predicted risk at the given time against the anticipated value of future information. To address this challenge, we analyze the optimal timing of early alerts using a Partially Observable Markov Decision Process (POMDP) with asymmetric reward. To learn an optimal alerting policy, we then propose a model-free reinforcement learning (RL) framework tailored to long-term clinical event surveillance from EHRs. Our proposed framework overcomes the pervasive issue of right-censoring in offline EHRs by leveraging a pseudo-label imputation approach. We also analytically demonstrate that entropy-regularized RL enables post-hoc threshold calibration to adapt the learned policy to specific preferences regarding the importance of earliness versus specificity without retraining. Systematic evaluations in synthetic data reveal that the advantage of RL-based look-ahead planning is maximized when diagnostic evidence emerges in predictable information bursts. Finally, real-world validations on two clinical cohorts show that our policy achieves an actionable lead time of 20.9 months prior to Alzheimer’s disease diagnosis and 8.7 months prior to autism diagnosis in a pediatric cohort, both at 90% specificity. Our method and results provide a generalizable blueprint for optimal surveillance of chronic disease processes.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-weng26a, title = {Censoring-Aware Reinforcement Learning to Optimize Early Risk Alerts from Longitudinal Clinical Data}, author = {Weng, Qin and Goldstein, Benjamin and Engelhard, Matthew M.}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {2151--2174}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/weng26a/weng26a.pdf}, url = {https://proceedings.mlr.press/v340/weng26a.html}, abstract = {Early recognition of chronic conditions is critical to ensure patients receive timely interventions and support. Passive surveillance of routine electronic health records (EHRs) provides information about longitudinal health trajectories that can support prompt recognition and inform associated early actions. However, relevant information is acquired at a different rate for each patient, and there is an inherent trade-off between the earliness versus the specificity of diagnosis and related actions. Therefore, determining when to alert providers about a likely chronic condition requires us to weigh the predicted risk at the given time against the anticipated value of future information. To address this challenge, we analyze the optimal timing of early alerts using a Partially Observable Markov Decision Process (POMDP) with asymmetric reward. To learn an optimal alerting policy, we then propose a model-free reinforcement learning (RL) framework tailored to long-term clinical event surveillance from EHRs. Our proposed framework overcomes the pervasive issue of right-censoring in offline EHRs by leveraging a pseudo-label imputation approach. We also analytically demonstrate that entropy-regularized RL enables post-hoc threshold calibration to adapt the learned policy to specific preferences regarding the importance of earliness versus specificity without retraining. Systematic evaluations in synthetic data reveal that the advantage of RL-based look-ahead planning is maximized when diagnostic evidence emerges in predictable information bursts. Finally, real-world validations on two clinical cohorts show that our policy achieves an actionable lead time of 20.9 months prior to Alzheimer’s disease diagnosis and 8.7 months prior to autism diagnosis in a pediatric cohort, both at 90% specificity. Our method and results provide a generalizable blueprint for optimal surveillance of chronic disease processes.} }
Endnote
%0 Conference Paper %T Censoring-Aware Reinforcement Learning to Optimize Early Risk Alerts from Longitudinal Clinical Data %A Qin Weng %A Benjamin Goldstein %A Matthew M. Engelhard %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-weng26a %I PMLR %P 2151--2174 %U https://proceedings.mlr.press/v340/weng26a.html %V 340 %X Early recognition of chronic conditions is critical to ensure patients receive timely interventions and support. Passive surveillance of routine electronic health records (EHRs) provides information about longitudinal health trajectories that can support prompt recognition and inform associated early actions. However, relevant information is acquired at a different rate for each patient, and there is an inherent trade-off between the earliness versus the specificity of diagnosis and related actions. Therefore, determining when to alert providers about a likely chronic condition requires us to weigh the predicted risk at the given time against the anticipated value of future information. To address this challenge, we analyze the optimal timing of early alerts using a Partially Observable Markov Decision Process (POMDP) with asymmetric reward. To learn an optimal alerting policy, we then propose a model-free reinforcement learning (RL) framework tailored to long-term clinical event surveillance from EHRs. Our proposed framework overcomes the pervasive issue of right-censoring in offline EHRs by leveraging a pseudo-label imputation approach. We also analytically demonstrate that entropy-regularized RL enables post-hoc threshold calibration to adapt the learned policy to specific preferences regarding the importance of earliness versus specificity without retraining. Systematic evaluations in synthetic data reveal that the advantage of RL-based look-ahead planning is maximized when diagnostic evidence emerges in predictable information bursts. Finally, real-world validations on two clinical cohorts show that our policy achieves an actionable lead time of 20.9 months prior to Alzheimer’s disease diagnosis and 8.7 months prior to autism diagnosis in a pediatric cohort, both at 90% specificity. Our method and results provide a generalizable blueprint for optimal surveillance of chronic disease processes.
APA
Weng, Q., Goldstein, B. & Engelhard, M.M.. (2026). Censoring-Aware Reinforcement Learning to Optimize Early Risk Alerts from Longitudinal Clinical Data. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:2151-2174 Available from https://proceedings.mlr.press/v340/weng26a.html.

Related Material