2014/02/09 by Arthur Guez, David Silver, Guez, Arthur +3 · 2 citations
Computer Science · Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #Artificial intelligence #Bayes' theorem #Bayesian inference #Bayesian probability #Bayesian programming #Computer science #FOS: Computer and information sciences #Frequentist inference #Inference #Leverage (statistics) #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Machine Learning and Algorithms #Machine learning #Mathematics #Parametric statistics #Probabilistic logic #Reinforcement Learning in Robotics #Sampling (signal processing) #Statistics #Variable-order Bayesian network #cs.AI #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1402.1958
published in arXiv (Cornell University) (Cornell University) · 11 pages, 11 figures
arxiv created 2014/02/09 · openalex publication_date 2014/02/09 · arxiv updated 2014/02/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The computational costs of inference and planning have confined Bayesian model-based reinforcement learning to one of two dismal fates: powerful Bayes-adaptive planning but only for simplistic models, or powerful, Bayesian non-parametric models but using simple, myopic planning strategies such as Thompson sampling. We ask whether it is feasible and truly beneficial to combine rich probabilistic models with a closer approximation to fully Bayesian planning. First, we use a collection of counterexamples to show formal problems with the over-optimism inherent in Thompson sampling. Then we leverage state-of-the-art techniques in efficient Bayes-adaptive planning and non-parametric Bayesian methods to perform qualitatively better than both existing conventional algorithms and Thompson sampling on two contextual bandit-like problems.