vix.ing · top · new · best · stats

Deep Exploration via Randomized Value Functions

2017/03/22 by Ian Osband, Benjamin Van Roy, Osband, Ian +5 · 31 citations
Computer Science · Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #Adversarial Robustness in Machine Learning #Reinforcement Learning in Robotics #cs.AI #cs.LG #stat.ML

paper · pdf · doi:10.48550/arxiv.1703.07608

Accepted for publication in Journal of Machine Learning Research 2019

arxiv created 2019/09/23 · arxiv updated 2019/09/25

Abstract

We study the use of randomized value functions to guide deep exploration in reinforcement learning. This offers an elegant means for synthesizing statistically and computationally efficient exploration with common practical approaches to value function learning. We present several reinforcement learning algorithms that leverage randomized value functions and demonstrate their efficacy through computational studies. We also prove a regret bound that establishes statistical efficiency with a tabular representation.

Citations

Cited by

Related