2020/06/17 by Hua Zheng, Zheng, Hua, Wei Xie +4
Biochemistry, Genetics and Molecular Biology · Computer Science · Mathematics · #Advanced Multi-Objective Optimization Algorithms #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Viral Infectious Diseases and Gene Expression in Insects #cs.AI #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2006.09919
12 pages, 1 figures. To appear in the Proceedings of the 2020 Winter Simulation Conference (WSC)
arxiv created 2020/06/17 · openalex publication_date 2020/06/17 · arxiv updated 2020/06/18 · openalex created_date 2020/06/25 · openalex updated_date 2026/07/28
Biopharmaceutical manufacturing faces critical challenges, including complexity, high variability, lengthy lead time, and limited historical data and knowledge of the underlying system stochastic process. To address these challenges, we propose a green simulation assisted model-based reinforcement learning to support process online learning and guide dynamic decision making. Basically, the process model risk is quantified by the posterior distribution. At any given policy, we predict the expected system response with prediction risk accounting for both inherent stochastic uncertainty and model risk. Then, we propose green simulation assisted reinforcement learning and derive the mixture proposal distribution of decision process and likelihood ratio based metamodel for the policy gradient, which can selectively reuse process trajectory outputs collected from previous experiments to increase the simulation data-efficiency, improve the policy gradient estimation accuracy, and speed up the search for the optimal policy. Our numerical study indicates that the proposed approach demonstrates the promising performance.