2019/06/26 by Onur Çelik, Celik, Onur, Hany Abdulsamad +3
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Autonomous Vehicle Technology and Safety #FOS: Electrical engineering #Reinforcement Learning in Robotics #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.1906.11003
openalex publication_date 2019/06/26 · openalex created_date 2022/07/22 · openalex updated_date 2026/07/28
Iterative trajectory optimization techniques for non-linear dynamical systems\nare among the most powerful and sample-efficient methods of model-based\nreinforcement learning and approximate optimal control. By leveraging\ntime-variant local linear-quadratic approximations of system dynamics and\nreward, such methods can find both a target-optimal trajectory and time-variant\noptimal feedback controllers. However, the local linear-quadratic assumptions\nare a major source of optimization bias that leads to catastrophic greedy\nupdates, raising the issue of proper regularization. Moreover, the approximate\nmodels' disregard for any physical state-action limits of the system causes\nfurther aggravation of the problem, as the optimization moves towards\nunreachable areas of the state-action space. In this paper, we address the\nissue of constrained systems in the scenario of online-fitted stochastic linear\ndynamics. We propose modeling state and action physical limits as probabilistic\nchance constraints linear in both state and action and introduce a new\ntrajectory optimization technique that integrates these probabilistic\nconstraints by optimizing a relaxed quadratic program. Our empirical\nevaluations show a significant improvement in learning robustness, which\nenables our approach to perform more effective updates and avoid premature\nconvergence observed in state-of-the-art algorithms.\n