2014/06/16 by Yasin Abbasi-Yadkori, Csaba Szepesvári, Abbasi-Yadkori, Yasin +1 · 3 citations
Computer Science · Decision Sciences · Engineering · Mathematics · #Advanced Bandit Algorithms Research #Advanced Control Systems Optimization #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Markov Chains and Monte Carlo Methods #Reinforcement Learning in Robotics
paper · pdf · doi:10.48550/arxiv.1406.3926
openalex publication_date 2014/06/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We study Bayesian optimal control of a general class of smoothly\nparameterized Markov decision problems. Since computing the optimal control is\ncomputationally expensive, we design an algorithm that trades off performance\nfor computational efficiency. The algorithm is a lazy posterior sampling method\nthat maintains a distribution over the unknown parameter. The algorithm changes\nits policy only when the variance of the distribution is reduced sufficiently.\nImportantly, we analyze the algorithm and show the precise nature of the\nperformance vs. computation tradeoff. Finally, we show the effectiveness of the\nmethod on a web server control application.\n