2019/01/24 by Chun‐Hao Chang, Chang, Chun-Hao, Mingjie Mai +3
Computer Science · Health Professions · Medicine · #FOS: Computer and information sciences #Healthcare Operations and Scheduling Optimization #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Healthcare #Sepsis Diagnosis and Treatment
paper · pdf · doi:10.48550/arxiv.1901.09699
openalex publication_date 2019/01/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Imagine a patient in critical condition. What and when should be measured to forecast detrimental events, especially under the budget constraints? We answer this question by deep reinforcement learning (RL) that jointly minimizes the measurement cost and maximizes predictive gain, by scheduling strategically-timed measurements. We learn our policy to be dynamically dependent on the patient's health history. To scale our framework to exponentially large action space, we distribute our reward in a sequential setting that makes the learning easier. In our simulation, our policy outperforms heuristic-based scheduling with higher predictive gain and lower cost. In a real-world ICU mortality prediction task (MIMIC3), our policies reduce the total number of measurements by 31% or improve predictive gain by a factor of 3 as compared to physicians, under the off-policy policy evaluation.