2012/05/08 by Marek Petrik, Petrik, Marek
Computer Science · Decision Sciences · Engineering · #Adaptive Dynamic Programming Control #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Reinforcement Learning in Robotics #Risk and Portfolio Optimization #Smart Grid Energy Management
paper · pdf · doi:10.48550/arxiv.1205.1782
openalex publication_date 2012/05/08 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Approximate dynamic programming is a popular method for solving large Markov\ndecision processes. This paper describes a new class of approximate dynamic\nprogramming (ADP) methods- distributionally robust ADP-that address the curse\nof dimensionality by minimizing a pessimistic bound on the policy loss. This\napproach turns ADP into an optimization problem, for which we derive new\nmathematical program formulations and analyze its properties. DRADP improves on\nthe theoretical guarantees of existing ADP methods-it guarantees convergence\nand L1 norm based error bounds. The empirical evaluation of DRADP shows that\nthe theoretical guarantees translate well into good performance on benchmark\nproblems.\n