Confounding Masquerading as Improvement: A Systematic Evaluation of Offline Reinforcement Learning for Stroke Antithrombotic Treatment in a 129,000-Patient Registry
2026-08-31 • Machine Learning
Machine Learning
AI summaryⓘ
The authors studied how well different offline reinforcement learning (RL) methods can suggest treatments for stroke patients compared to doctors. They found that some initial measures showed small improvements, but these were partly due to a hidden problem where the reward signals also captured patients' starting health severity, not just treatment effects. After adjusting for this confounding, the apparent benefits mostly disappeared. Their analyses suggest no clear overall advantage of the RL policies in improving patient outcomes, and they propose a checklist to more carefully evaluate such models in the future.
offline reinforcement learningFitted Q-Evaluationreward confoundingacute ischemic strokeEarly Neurological Deteriorationcausal inferenceDeep Mutual Learning (DML)modified Rankin Scale (mRS)gradient boosting machines (GBM)policy evaluation
Authors
Kihun Rhee
Abstract
Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL algorithm families and 14 reward designs in 44,894 post-2018 acute ischemic stroke patients from a nationwide registry (N = 129,033). Standard Fitted Q-Evaluation (FQE) yields an apparent policy-improvement estimate of +0.0069; adding an Early Neurological Deterioration penalty increases it to +0.0101. We identify reward-embedded confounding, in which a proxy terminal reward encodes baseline severity and prognosis as well as treatment efficacy. A 2 x 2 factorial analysis finds that terminal reward confounding accounts for 218.6% of the observed signal change, so its removal overshoots the null. After DML-inspired GBM reward residualization, the FQE estimate attenuates to +0.0033 (p = 0.132), and full deconfounding yields +0.0025 (p = 0.291). FQE-based diagnostics, T-learner analyses, and direct recurrence analyses converge away from a clinically meaningful aggregate improvement. A 1-year mRS factorial analysis replicates the attenuation. We provide an empirically motivated six-step evaluation checklist. NIHSS-stratified heterogeneity is hypothesis-generating for prospective trial design; hospital-level disagreement does not persist after full reward deconfounding.