2018/08/06 by Dror Freirich, Freirich, Dror, Ron Meir +3 · 3 citations
Computer Science · Decision Sciences · Economics, Econometrics and Finance · Physics and Astronomy · #FOS: Computer and information sciences #Forecasting Techniques and Applications #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Model Reduction and Neural Networks #Reinforcement Learning in Robotics #Sports Analytics and Performance
paper · pdf · doi:10.48550/arxiv.1808.01960
openalex publication_date 2018/08/06 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The recently proposed distributional approach to reinforcement learning\n(DiRL) is centered on learning the distribution of the reward-to-go, often\nreferred to as the value distribution. In this work, we show that the\ndistributional Bellman equation, which drives DiRL methods, is equivalent to a\ngenerative adversarial network (GAN) model. In this formulation, DiRL can be\nseen as learning a deep generative model of the value distribution, driven by\nthe discrepancy between the distribution of the current value, and the\ndistribution of the sum of current reward and next value. We use this insight\nto propose a GAN-based approach to DiRL, which leverages the strengths of GANs\nin learning distributions of high-dimensional data. In particular, we show that\nour GAN approach can be used for DiRL with multivariate rewards, an important\nsetting which cannot be tackled with prior methods. The multivariate setting\nalso allows us to unify learning the distribution of values and state\ntransitions, and we exploit this idea to devise a novel exploration method that\nis driven by the discrepancy in estimating both values and states.\n