vix.ing · top · new · best · stats · spec

Mean-Variance Policy Iteration for Risk-Averse Reinforcement Learning

2020/04/22 by Shangtong Zhang, Bo Liu, Zhang, Shangtong +3 · 4 citations
Computer Science · Medicine · #Adversarial Robustness in Machine Learning #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Neurological disorders and treatments #Reinforcement Learning in Robotics

paper · pdf · doi:10.48550/arxiv.2004.10888

openalex publication_date 2020/04/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/01

Abstract

We present a mean-variance policy iteration (MVPI) framework for risk-averse control in a discounted infinite horizon MDP optimizing the variance of a per-step reward random variable. MVPI enjoys great flexibility in that any policy evaluation method and risk-neutral control method can be dropped in for risk-averse control off the shelf, in both on- and off-policy settings. This flexibility reduces the gap between risk-neutral control and risk-averse control and is achieved by working on a novel augmented MDP directly. We propose risk-averse TD3 as an example instantiating MVPI, which outperforms vanilla TD3 and many previous risk-averse control methods in challenging Mujoco robot simulation tasks under a risk-aware performance metric. This risk-averse TD3 is the first to introduce deterministic policies and off-policy learning into risk-averse reinforcement learning, both of which are key to the performance boost we show in Mujoco domains.

Citations

Cited by

Related