vix.ing · top · new · best · stats · spec

Stochastic Approximation for Risk-aware Markov Decision Processes

2018/05/11 by Wenjie Huang, William B. Haskell, Huang, Wenjie +1 · 1 citation
Decision Sciences · Mathematics · #Advanced Bandit Algorithms Research #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #FOS: Mathematics #Optimization and Control (math.OC) #Risk and Portfolio Optimization #Statistical Methods and Inference

paper · pdf · doi:10.48550/arxiv.1805.04238

openalex publication_date 2018/05/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We develop a stochastic approximation-type algorithm to solve finite state/action, infinite-horizon, risk-aware Markov decision processes. Our algorithm has two loops. The inner loop computes the risk by solving a stochastic saddle-point problem. The outer loop performs Q-learning to compute an optimal risk-aware policy. Several widely investigated risk measures (e.g. conditional value-at-risk, optimized certainty equivalent, and absolute semi-deviation) are covered by our algorithm. Almost sure convergence and the convergence rate of the algorithm are established. For an error tolerance ε>0 for the optimal Q-value estimation gap and learning rate k∈(1/2, 1], the overall convergence rate of our algorithm is Ω((ln(1/δε)/ε2)1/k+(ln(1/ε))1/(1-k)) with probability at least 1-δ.

Citations

Cited by

Related