2022/02/09 by Elynn Chen, Chen, Elynn, Li, Sai +1 · 1 citation
Computer Science · Decision Sciences · #Advanced Bandit Algorithms Research #FOS: Computer and information sciences #Gaussian Processes and Bayesian Inference #Machine Learning (cs.LG) #Methodology (stat.ME) #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.2202.04709
openalex publication_date 2022/02/09 · openalex created_date 2022/05/05 · openalex updated_date 2026/07/28
Time-inhomogeneous finite-horizon Markov decision processes (MDP) are frequently employed to model decision-making in dynamic treatment regimes and other statistical reinforcement learning (RL) scenarios. These fields, especially healthcare and business, often face challenges such as high-dimensional state spaces and time-inhomogeneity of the MDP process, compounded by insufficient sample availability which complicates informed decision-making. To overcome these challenges, we investigate knowledge transfer within time-inhomogeneous finite-horizon MDP by leveraging data from both a target RL task and several related source tasks. We have developed transfer learning (TL) algorithms that are adaptable for both batch and online Q-learning, integrating valuable insights from offline source studies. The proposed transfer Q-learning algorithm contains a novel \em re-targeting step that enables \em cross-stage transfer along multiple stages in an RL task, besides the usual \em cross-task transfer for supervised learning. We establish the first theoretical justifications of TL in RL tasks by showing a faster rate of convergence of the Q^*-function estimation in the offline RL transfer, and a lower regret bound in the offline-to-online RL transfer under stage-wise reward similarity and mild design similarity across tasks. Empirical evidence from both synthetic and real datasets is presented to evaluate the proposed algorithm and support our theoretical results.