[edit]
Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry
Proceedings of the 11th Machine Learning for Healthcare Conference, PMLR 340:1689-1725, 2026.
Abstract
Recent offline reinforcement learning (RL) studies report data-driven policies outperform physician decisions by 10–32% on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs on 44,894 post-2018 acute ischemic stroke patients from the nationwide stroke registry ($N = 129{,}033$). Standard Fitted Q-Evaluation yields $V(\mathrm{imp}) = +0.0069$ ($p = 0.048$); adding an Early Neurological Deterioration penalty strengthens the signal to $+0.0101$ ($p = 0.0002$, approximate E-value $= 18.37$)—estimates that could support a premature positive policy-improvement claim despite relying on a confounded reward function. We identify reward-embedded confounding, where the proxy terminal reward encodes baseline severity/prognosis information in addition to any treatment efficacy signal. E-value analysis does not evaluate this pathway because the confounding is carried through the reward channel. A $2 \times 2$ factorial experiment reveals that terminal reward confounding alone more than eliminates the apparent improvement, accounting for 218.6% of the $r_{\mathrm{END}} \to r_{\mathrm{full}}$ signal change (i.e., removing terminal confounding overshoots null). After DML-inspired GBM reward residualization, $V(\mathrm{imp})$ attenuates to $+0.0033$ ($p = 0.132$); full deconfounding yields $+0.0025$ ($p = 0.291$). Three independent approaches converge away from a clinically meaningful aggregate policy-improvement claim: FQE-based reward deconfounding, a T-learner showing 75–90% attenuation with a clinically small residual ($<6%$ MCID), and direct ischemic-stroke recurrence analysis (IPW/AIPW). A within-registry 1- year mRS factorial ($N = 35{,}744$) replicates the pattern: $V(\mathrm{imp}) = +0.0133$ attenuates to $-0.0004$ after reward residualization (103.3% attenuation), indicating that the finding is not specific to the 3-month mRS horizon in this registry. We present an empirically motivated six-step evaluation checklist; applied retrospectively, three highly cited clinical RL papers do not report the FQE confirmation specified by Step 1. Despite the overall null, National Institutes of Health Stroke Scale (NIHSS)-stratified heterogeneity (T-learner conditional average treatment effect (CATE) ratio $4.6\times$, corroborated by Q-evaluation $3.4\times$ gradient) identifies a hypothesis-generating subgroup for prospective trial design; hospital-level disagreement did not persist after full reward deconfounding.