Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry

Kihun Rhee
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1689-1725, 2026.

Abstract

Recent offline reinforcement learning (RL) studies report data-driven policies outperform physician decisions by 10–32% on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs on 44,894 post-2018 acute ischemic stroke patients from the nationwide stroke registry ($N = 129{,}033$). Standard Fitted Q-Evaluation yields $V(\mathrm{imp}) = +0.0069$ ($p = 0.048$); adding an Early Neurological Deterioration penalty strengthens the signal to $+0.0101$ ($p = 0.0002$, approximate E-value $= 18.37$)—estimates that could support a premature positive policy-improvement claim despite relying on a confounded reward function. We identify reward-embedded confounding, where the proxy terminal reward encodes baseline severity/prognosis information in addition to any treatment efficacy signal. E-value analysis does not evaluate this pathway because the confounding is carried through the reward channel. A $2 \times 2$ factorial experiment reveals that terminal reward confounding alone more than eliminates the apparent improvement, accounting for 218.6% of the $r_{\mathrm{END}} \to r_{\mathrm{full}}$ signal change (i.e., removing terminal confounding overshoots null). After DML-inspired GBM reward residualization, $V(\mathrm{imp})$ attenuates to $+0.0033$ ($p = 0.132$); full deconfounding yields $+0.0025$ ($p = 0.291$). Three independent approaches converge away from a clinically meaningful aggregate policy-improvement claim: FQE-based reward deconfounding, a T-learner showing 75–90% attenuation with a clinically small residual ($<6%$ MCID), and direct ischemic-stroke recurrence analysis (IPW/AIPW). A within-registry 1- year mRS factorial ($N = 35{,}744$) replicates the pattern: $V(\mathrm{imp}) = +0.0133$ attenuates to $-0.0004$ after reward residualization (103.3% attenuation), indicating that the finding is not specific to the 3-month mRS horizon in this registry. We present an empirically motivated six-step evaluation checklist; applied retrospectively, three highly cited clinical RL papers do not report the FQE confirmation specified by Step 1. Despite the overall null, National Institutes of Health Stroke Scale (NIHSS)-stratified heterogeneity (T-learner conditional average treatment effect (CATE) ratio $4.6\times$, corroborated by Q-evaluation $3.4\times$ gradient) identifies a hypothesis-generating subgroup for prospective trial design; hospital-level disagreement did not persist after full reward deconfounding.

Cite this Paper


BibTeX
@InProceedings{pmlr-v340-rhee26a, title = {Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry}, author = {Rhee, Kihun}, booktitle = {Proceedings of the 11th Machine Learning for Healthcare Conference}, pages = {1689--1725}, year = {2026}, editor = {Krishnan, Rahul G. and van Amsterdam, Wouter A. C. and Chopra, Sumit and Overgaard, Shauna and Hughes, Michael and Ötleş, Erkin and Shen, Yiqiu and Shanmugam, Divya and Nayan, Madhur and Engelhard, Matthew and Fackler, Jim and Oberst, Michael}, volume = {340}, series = {Proceedings of Machine Learning Research}, month = {12--14 Aug}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v340/main/assets/rhee26a/rhee26a.pdf}, url = {https://proceedings.mlr.press/v340/rhee26a.html}, abstract = {Recent offline reinforcement learning (RL) studies report data-driven policies outperform physician decisions by 10–32% on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs on 44,894 post-2018 acute ischemic stroke patients from the nationwide stroke registry ($N = 129{,}033$). Standard Fitted Q-Evaluation yields $V(\mathrm{imp}) = +0.0069$ ($p = 0.048$); adding an Early Neurological Deterioration penalty strengthens the signal to $+0.0101$ ($p = 0.0002$, approximate E-value $= 18.37$)—estimates that could support a premature positive policy-improvement claim despite relying on a confounded reward function. We identify reward-embedded confounding, where the proxy terminal reward encodes baseline severity/prognosis information in addition to any treatment efficacy signal. E-value analysis does not evaluate this pathway because the confounding is carried through the reward channel. A $2 \times 2$ factorial experiment reveals that terminal reward confounding alone more than eliminates the apparent improvement, accounting for 218.6% of the $r_{\mathrm{END}} \to r_{\mathrm{full}}$ signal change (i.e., removing terminal confounding overshoots null). After DML-inspired GBM reward residualization, $V(\mathrm{imp})$ attenuates to $+0.0033$ ($p = 0.132$); full deconfounding yields $+0.0025$ ($p = 0.291$). Three independent approaches converge away from a clinically meaningful aggregate policy-improvement claim: FQE-based reward deconfounding, a T-learner showing 75–90% attenuation with a clinically small residual ($<6%$ MCID), and direct ischemic-stroke recurrence analysis (IPW/AIPW). A within-registry 1- year mRS factorial ($N = 35{,}744$) replicates the pattern: $V(\mathrm{imp}) = +0.0133$ attenuates to $-0.0004$ after reward residualization (103.3% attenuation), indicating that the finding is not specific to the 3-month mRS horizon in this registry. We present an empirically motivated six-step evaluation checklist; applied retrospectively, three highly cited clinical RL papers do not report the FQE confirmation specified by Step 1. Despite the overall null, National Institutes of Health Stroke Scale (NIHSS)-stratified heterogeneity (T-learner conditional average treatment effect (CATE) ratio $4.6\times$, corroborated by Q-evaluation $3.4\times$ gradient) identifies a hypothesis-generating subgroup for prospective trial design; hospital-level disagreement did not persist after full reward deconfounding.} }
Endnote
%0 Conference Paper %T Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry %A Kihun Rhee %B Proceedings of the 11th Machine Learning for Healthcare Conference %C Proceedings of Machine Learning Research %D 2026 %E Rahul G. Krishnan %E Wouter A. C. van Amsterdam %E Sumit Chopra %E Shauna Overgaard %E Michael Hughes %E Erkin Ötleş %E Yiqiu Shen %E Divya Shanmugam %E Madhur Nayan %E Matthew Engelhard %E Jim Fackler %E Michael Oberst %F pmlr-v340-rhee26a %I PMLR %P 1689--1725 %U https://proceedings.mlr.press/v340/rhee26a.html %V 340 %X Recent offline reinforcement learning (RL) studies report data-driven policies outperform physician decisions by 10–32% on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs on 44,894 post-2018 acute ischemic stroke patients from the nationwide stroke registry ($N = 129{,}033$). Standard Fitted Q-Evaluation yields $V(\mathrm{imp}) = +0.0069$ ($p = 0.048$); adding an Early Neurological Deterioration penalty strengthens the signal to $+0.0101$ ($p = 0.0002$, approximate E-value $= 18.37$)—estimates that could support a premature positive policy-improvement claim despite relying on a confounded reward function. We identify reward-embedded confounding, where the proxy terminal reward encodes baseline severity/prognosis information in addition to any treatment efficacy signal. E-value analysis does not evaluate this pathway because the confounding is carried through the reward channel. A $2 \times 2$ factorial experiment reveals that terminal reward confounding alone more than eliminates the apparent improvement, accounting for 218.6% of the $r_{\mathrm{END}} \to r_{\mathrm{full}}$ signal change (i.e., removing terminal confounding overshoots null). After DML-inspired GBM reward residualization, $V(\mathrm{imp})$ attenuates to $+0.0033$ ($p = 0.132$); full deconfounding yields $+0.0025$ ($p = 0.291$). Three independent approaches converge away from a clinically meaningful aggregate policy-improvement claim: FQE-based reward deconfounding, a T-learner showing 75–90% attenuation with a clinically small residual ($<6%$ MCID), and direct ischemic-stroke recurrence analysis (IPW/AIPW). A within-registry 1- year mRS factorial ($N = 35{,}744$) replicates the pattern: $V(\mathrm{imp}) = +0.0133$ attenuates to $-0.0004$ after reward residualization (103.3% attenuation), indicating that the finding is not specific to the 3-month mRS horizon in this registry. We present an empirically motivated six-step evaluation checklist; applied retrospectively, three highly cited clinical RL papers do not report the FQE confirmation specified by Step 1. Despite the overall null, National Institutes of Health Stroke Scale (NIHSS)-stratified heterogeneity (T-learner conditional average treatment effect (CATE) ratio $4.6\times$, corroborated by Q-evaluation $3.4\times$ gradient) identifies a hypothesis-generating subgroup for prospective trial design; hospital-level disagreement did not persist after full reward deconfounding.
APA
Rhee, K.. (2026). Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry. Proceedings of the 11th Machine Learning for Healthcare Conference, in Proceedings of Machine Learning Research 340:1689-1725 Available from https://proceedings.mlr.press/v340/rhee26a.html.

Related Material