2019/09/23 by Yuan, Lim Zun, Hasanbeig, Mohammadhosein, Abate, Alessandro +1
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #FOS: Electrical engineering #Logic in Computer Science (cs.LO) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.1909.11591
We propose an actor-critic, model-free, and online Reinforcement Learning (RL) framework for continuous-state continuous-action Markov Decision Processes (MDPs) when the reward is highly sparse but encompasses a high-level temporal structure. We represent this temporal structure by a finite-state machine and construct an on-the-fly synchronised product with the MDP and the finite machine. The temporal structure acts as a guide for the RL agent within the product, where a modular Deep Deterministic Policy Gradient (DDPG) architecture is proposed to generate a low-level control policy. We evaluate our framework in a Mars rover experiment and we present the success rate of the synthesised policy.