vix.ing · top · new · best · stats · spec

Temporal Difference Learning with Continuous Time and State in the Stochastic Setting

2022/02/16 by Ziad Kobeissi, Francis Bach, Kobeissi, Ziad +1 · 1 citation
Computer Science · Decision Sciences · Economics, Econometrics and Finance · Energy · #Analysis of PDEs (math.AP) #Artificial Intelligence (cs.AI) #Climate Change Policy and Economics #Economic Policies and Impacts #Energy Efficiency and Management #FOS: Computer and information sciences #FOS: Mathematics #Machine Learning (cs.LG) #Optimization and Control (math.OC) #Reinforcement Learning in Robotics #Simulation Techniques and Applications

paper · pdf · doi:10.48550/arxiv.2202.07960

openalex publication_date 2022/02/16 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We consider the problem of continuous-time policy evaluation. This consists in learning through observations the value function associated with an uncontrolled continuous-time stochastic dynamic and a reward function. We propose two original variants of the well-known TD(0) method using vanishing time steps. One is model-free and the other is model-based. For both methods, we prove theoretical convergence rates that we subsequently verify through numerical simulations. Alternatively, those methods can be interpreted as novel reinforcement learning approaches for approximating solutions of linear PDEs (partial differential equations) or linear BSDEs (backward stochastic differential equations).

Cited by

Related