vix.ing · top · new · best · stats

Model-Based Reinforcement Learning via Meta-Policy Optimization

2018/09/14 by Ignasi Clavera, Jonas Rothfuss, Clavera, Ignasi +9 · 15 citations
Computer Science · Mathematics · #Artificial Intelligence (cs.AI) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Data Classification #Reinforcement Learning in Robotics #cs.AI #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.1809.05214

First 2 authors contributed equally. Accepted for Conference on Robot Learning (CoRL)

arxiv created 2018/09/14 · openalex publication_date 2018/09/14 · arxiv updated 2018/09/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Model-based reinforcement learning approaches carry the promise of being data efficient. However, due to challenges in learning dynamics models that sufficiently match the real-world dynamics, they struggle to achieve the same asymptotic performance as model-free methods. We propose Model-Based Meta-Policy-Optimization (MB-MPO), an approach that foregoes the strong reliance on accurate learned dynamics models. Using an ensemble of learned dynamic models, MB-MPO meta-learns a policy that can quickly adapt to any model in the ensemble with one policy gradient step. This steers the meta-policy towards internalizing consistent dynamics predictions among the ensemble while shifting the burden of behaving optimally w.r.t. the model discrepancies towards the adaptation step. Our experiments show that MB-MPO is more robust to model imperfections than previous model-based approaches. Finally, we demonstrate that our approach is able to match the asymptotic performance of model-free methods while requiring significantly less experience.

Cited by

Related