Exploring Time-Step Size in Reinforcement Learning for Sepsis Treatment Open Access

Sun, Yingchuan (Spring 2026)

Permanent URL: https://etd.library.emory.edu/concern/etds/5h73px592?locale=en
Published

Abstract

Existing studies on reinforcement learning (RL) for sepsis management have mostly followed an established problem setup, in which patient data are aggregated into 4-hour time steps. Although concerns have been raised regarding the coarseness of this time-step size, which might distort patient dynamics and lead to suboptimal treatment policies, the extent to which this is a problem in practice remains unex- plored. In this work, we conducted empirical experiments for a controlled comparison of four time-step sizes (∆t = 1, 2, 4, 8 h) on this domain, following an identical offline RL pipeline. To enable a fair comparison across time-step sizes, we designed action re-mapping methods that allow for evaluation of policies on datasets with different time-step sizes, and conducted cross-∆t model selections under two policy learning setups. Our goal was to quantify how time-step size influences state representation learning, behavior cloning, policy training, and off-policy evaluation. Our results show that performance trends across ∆t vary as learning setups change, while poli- cies learned at finer time-step sizes (∆t = 1 h and 2 h) using a static behavior policy achieve the overall best performance and stability. Our work highlights time-step size as a core design choice in offline RL for healthcare and provides evidence supporting alternatives beyond the conventional 4-hour setup.

Table of Contents

1 Introduction 1

2 Related Work 4

3 Background & Problem Setup 7

3.1 Time Step Discretization 7

3.2 Offline RL Objective 7

4 Experimental Setup 9

4.1 Cohort Construction and Preprocessing 10

4.2 State Representation Learning 11

4.3 Behavior Cloning 12

4.4 Policy Learning 13

4.5 Policy Evaluation & Selection 13

4.6 Policy Analysis 17

5 Results 18

5.1 Cohort Statistics 18

5.2 State Representation Learning 19

5.3 Behavior Cloning 19

5.4 Policy Learning, Evaluation & Selection 20

5.5 Test Performance 22

6 Conclusion & Discussion 24

A Extended Methods 27

B Extended Results 29

Bibliography 33

About this Master's Thesis

Rights statement
  • Permission granted by the author to include this thesis or dissertation in this repository. All rights reserved by the author. Please contact the author for information regarding the reproduction and use of this thesis or dissertation.
School
Department
Degree
Submission
Language
  • English
Research Field
Keyword
Committee Chair / Thesis Advisor
Committee Members
Last modified

Primary PDF

Supplemental Files