vix.ing · top · new · best · stats · spec

Policy iteration using Q-functions: Linear dynamics with multiplicative noise

2022/12/02 by Peter Coppens, Panagiotis Patrinos, Coppens, Peter +1 · 1 citation
Computer Science · #FOS: Mathematics #Optimization and Control (math.OC) #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2212.01192

openalex publication_date 2022/12/02 · openalex created_date 2022/12/17 · openalex updated_date 2026/07/28

Abstract

This paper presents a novel model-free and fully data-driven policy iteration scheme for quadratic regulation of linear dynamics with state- and input-multiplicative noise. The implementation is similar to the least-squares temporal difference scheme for Markov decision processes, estimating Q-functions by solving a least-squares problem with instrumental variables. The scheme is compared with a model-based system identification scheme and natural policy gradient through numerical experiments.

Cited by

Related