Back

Offline Reinforcement Learning for Out-of-Distribution ICU Sepsis Decision Support

Arasteh, E.; Mirian, M. S.; Tavakol, M.

2026-08-25 health informatics
10.64898/2026.08.22.26361090 medRxiv
Show abstract

Offline reinforcement learning (RL) provides a promising framework for learning and evaluating treatment policies from logged clinical data, particularly in sequential decision-making settings where prospective exploration would be unsafe. In ICU sepsis management, however, it remains unclear whether offline RL policies retain stable behavior under increasingly severe out-of-distribution (OOD) patient cohorts. In this paper, we evaluate standard offline RL methods on three severity-enriched OOD test mixtures from the MIMIC-III benchmark dataset to determine whether offline policies retain a stable, actionsensitive decision-support signal. Under the shared learned-dynamics offpolicy evaluation (OPE) protocol, as the severe-OOD ratio increases from 25% to 75%, observed clinical survival declines from 67% to 49%, while the best offline method in each mixture receives model-predicted terminal survival values of 87%, 86%, and 85%, respectively. Because observed clinical survival and model-predicted terminal survival are different quantities, this contrast suggests a stable model-based decision-support signal under severity shift. We further present a secondary physiological stabilization analysis using an episode-level physiological stabilization score (EPSS), a heuristic summary of whether selected physiological variables move in favorable directions during follow-up. In this analysis, model-generated rollouts under offline policies receive higher EPSS values than matched logged clinical trajectories for several physiological components. Together, these results support learned-dynamics OPE as a useful severity-OOD stress test for offline RL policies in ICU sepsis, while leaving prospective and causal validation as necessary next steps.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.