vix.ing · top · new · best · stats

QMDP-Net: Deep Learning for Planning under Partial Observability

2017/03/20 by Peter Karkus, Péter Karkus, Karkus, Peter +4 · 5 citations
Computer Science · Mathematics · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Neural and Evolutionary Computing (cs.NE) #Reinforcement Learning in Robotics #cs.AI #cs.LG #cs.NE #stat.ML

paper · pdf · doi:10.48550/arxiv.1703.06692

NIPS 2017 camera-ready

openalex publication_date 2017/03/20 · arxiv created 2017/11/03 · arxiv updated 2017/11/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This paper introduces the QMDP-net, a neural network architecture for planning under partial observability. The QMDP-net combines the strengths of model-free learning and model-based planning. It is a recurrent policy network, but it represents a policy for a parameterized set of tasks by connecting a model with a planning algorithm that solves the model, thus embedding the solution structure of planning in a network learning architecture. The QMDP-net is fully differentiable and allows for end-to-end training. We train a QMDP-net on different tasks so that it can generalize to new ones in the parameterized task set and "transfer" to other similar tasks beyond the set. In preliminary experiments, QMDP-net showed strong performance on several robotic tasks in simulation. Interestingly, while QMDP-net encodes the QMDP algorithm, it sometimes outperforms the QMDP algorithm in the experiments, as a result of end-to-end learning.

Citations

Cited by

Related