2021/02/16 by Mohammadi Zaki, Zaki, Mohammadi, Avinash Mohan +5 · 2 citations
Computer Science · Decision Sciences · Engineering · #Advanced Bandit Algorithms Research #Age of Information Optimization #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Systems and Control (eess.SY) #cs.LG #cs.SY #eess.SY #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2102.08201
openalex publication_date 2021/02/16 · arxiv created 2021/07/03 · arxiv updated 2021/07/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We consider an improper reinforcement learning setting where a learner is given M base controllers for an unknown Markov decision process, and wishes to combine them optimally to produce a potentially new controller that can outperform each of the base ones. This can be useful in tuning across controllers, learnt possibly in mismatched or simulated environments, to obtain a good controller for a given target environment with relatively few trials. \par We propose a gradient-based approach that operates over a class of improper mixtures of the controllers. We derive convergence rate guarantees for the approach assuming access to a gradient oracle. The value function of the mixture and its gradient may not be available in closed-form; however, we show that we can employ rollouts and simultaneous perturbation stochastic approximation (SPSA) for explicit gradient descent optimization. Numerical results on (i) the standard control theoretic benchmark of stabilizing an inverted pendulum and (ii) a constrained queueing task show that our improper policy optimization algorithm can stabilize the system even when the base policies at its disposal are unstable\footnoteUnder review. Please do not distribute..