vix.ing · top · new · best · stats

Learning Continuous Control Policies by Stochastic Value Gradients

2015/10/30 by Nicolas Heess, Greg Wayne, Heess, Nicolas +9 · 112 citations
Computer Science · Engineering · Mathematics · #Advanced Control Systems Optimization #Artificial intelligence #Computer science #Control (management) #Economics #FOS: Computer and information sciences #Fault Detection and Control Systems #Machine Learning (cs.LG) #Mathematical economics #Mathematical optimization #Mathematics #Neural and Evolutionary Computing (cs.NE) #Optimal control #Reinforcement Learning in Robotics #Statistics #Stochastic control #Value (mathematics) #cs.LG #cs.NE

paper · pdf · doi:10.48550/arxiv.1510.09142

published in arXiv (Cornell University) (Cornell University) · 13 pages, NIPS 2015

arxiv created 2015/10/30 · openalex publication_date 2015/10/30 · arxiv updated 2015/11/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We present a unified framework for learning continuous control policies using backpropagation. It supports stochastic control by treating stochasticity in the Bellman equation as a deterministic function of exogenous noise. The product is a spectrum of general policy gradient algorithms that range from model-free methods with value functions to model-based methods without value functions. We use learned models but only require observations from the environment in- stead of observations from model-predicted trajectories, minimizing the impact of compounded model errors. We apply these algorithms first to a toy stochastic control problem and then to several physics-based control problems in simulation. One of these variants, SVG(1), shows the effectiveness of learning models, value functions, and policies simultaneously in continuous domains.

Citations

Cited by

Related