2015/12/02 by Yu, Pengqian, Yu, Jia Yuan, Xu, Huan
#FOS: Electrical engineering #FOS: Mathematics #Optimization and Control (math.OC) #Systems and Control (eess.SY) #electronic engineering #information engineering
paper · doi:10.48550/arxiv.1512.00583
Whereas classical Markov decision processes maximize the expected reward, we consider minimizing the risk. We propose to evaluate the risk associated to a given policy over a long-enough time horizon with the help of a central limit theorem. The proposed approach works whether the transition probabilities are known or not. We also provide a gradient-based policy improvement algorithm that converges to a local optimum of the risk objective.