vix.ing · top · new · best · stats · spec

Loaded DiCE: Trading off Bias and Variance in Any-Order Score Function\n Estimators for Reinforcement Learning

2019/09/23 by Gregory Farquhar, Shimon Whiteson, Farquhar, Gregory +3 · 1 citation
Computer Science · Psychology · #Advanced Multi-Objective Optimization Algorithms #Data Stream Mining Techniques #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Mental Health Research Topics #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.1909.10549

openalex publication_date 2019/09/23 · openalex created_date 2022/09/13 · openalex updated_date 2026/07/28

Abstract

Gradient-based methods for optimisation of objectives in stochastic settings\nwith unknown or intractable dynamics require estimators of derivatives. We\nderive an objective that, under automatic differentiation, produces\nlow-variance unbiased estimators of derivatives at any order. Our objective is\ncompatible with arbitrary advantage estimators, which allows the control of the\nbias and variance of any-order derivatives when using function approximation.\nFurthermore, we propose a method to trade off bias and variance of higher order\nderivatives by discounting the impact of more distant causal dependencies. We\ndemonstrate the correctness and utility of our objective in analytically\ntractable MDPs and in meta-reinforcement-learning for continuous control.\n

Citations

Cited by

Related