2012/11/26 by Sumeetpal S. Singh, Nicolas Chopin, ChopinNicolas +4
Computer Science · Engineering · Mathematics · #Control Systems and Identification #Gaussian Processes and Bayesian Inference #Reinforcement Learning in Robotics #cs.LG #stat.CO #stat.ML
paper · pdf · doi:10.48550/arxiv.1211.5901
arxiv created 2012/11/26 · arxiv updated 2012/11/27
We consider the inverse reinforcement learning problem, that is, the problem of learning from, and then predicting or mimicking a controller based on state/action data. We propose a statistical model for such data, derived from the structure of a Markov decision process. Adopting a Bayesian approach to inference, we show how latent variables of the model can be estimated, and how predictions about actions can be made, in a unified framework. A new Markov chain Monte Carlo (MCMC) sampler is devised for simulation from the posterior distribution. This step includes a parameter expansion step, which is shown to be essential for good convergence properties of the MCMC sampler. As an illustration, the method is applied to learning a human controller.