2020/01/13 by Joseph Futoma, Michael C. Hughes, Futoma, Joseph +3 · 2 citations
Computer Science · Health Professions · Medicine · #Chronic Disease Management Strategies #FOS: Computer and information sciences #Healthcare Operations and Scheduling Optimization #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning in Healthcare
paper · pdf · doi:10.48550/arxiv.2001.04032
openalex publication_date 2020/01/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Many medical decision-making tasks can be framed as partially observed Markov decision processes (POMDPs). However, prevailing two-stage approaches that first learn a POMDP and then solve it often fail because the model that best fits the data may not be well suited for planning. We introduce a new optimization objective that (a) produces both high-performing policies and high-quality generative models, even when some observations are irrelevant for planning, and (b) does so in batch off-policy settings that are typical in healthcare, when only retrospective data is available. We demonstrate our approach on synthetic examples and a challenging medical decision-making problem.