vix.ing · top · new · best · stats · spec

TD or not TD: Analyzing the Role of Temporal Differencing in Deep\n Reinforcement Learning

2018/06/04 by Artemij Amiranashvili, Amiranashvili, Artemij, Alexey Dosovitskiy +5
Computer Science · Neuroscience · Physics and Astronomy · #Artificial Intelligence (cs.AI) #Evolutionary Algorithms and Applications #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neural dynamics and brain function #stochastic dynamics and bifurcation

paper · pdf · doi:10.48550/arxiv.1806.01175

openalex publication_date 2018/06/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Our understanding of reinforcement learning (RL) has been shaped by\ntheoretical and empirical results that were obtained decades ago using tabular\nrepresentations and linear function approximators. These results suggest that\nRL methods that use temporal differencing (TD) are superior to direct Monte\nCarlo estimation (MC). How do these results hold up in deep RL, which deals\nwith perceptually complex environments and deep nonlinear models? In this\npaper, we re-examine the role of TD in modern deep RL, using specially designed\nenvironments that control for specific factors that affect performance, such as\nreward sparsity, reward delay, and the perceptual complexity of the task. When\ncomparing TD with infinite-horizon MC, we are able to reproduce classic results\nin modern settings. Yet we also find that finite-horizon MC is not inferior to\nTD, even when rewards are sparse or delayed. This makes MC a viable alternative\nto TD in deep RL.\n

Citations

Related