2016/05/31 by Ezio Bartocci, Bartocci, Ezio, Luca Bortolussi +7 · 1 citation
Computer Science · #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.1605.09703
Continuous-time Markov decision processes are an important class of models in\na wide range of applications, ranging from cyber-physical systems to synthetic\nbiology. A central problem is how to devise a policy to control the system in\norder to maximise the probability of satisfying a set of temporal logic\nspecifications. Here we present a novel approach based on statistical model\nchecking and an unbiased estimation of a functional gradient in the space of\npossible policies. The statistical approach has several advantages over\nconventional approaches based on uniformisation, as it can also be applied when\nthe model is replaced by a black box, and does not suffer from state-space\nexplosion. The use of a stochastic gradient to guide our search considerably\nimproves the efficiency of learning policies. We demonstrate the method on a\nproof-of-principle non-linear population model, showing strong performance in a\nnon-trivial task.\n