2019/09/15 by Hongwei Mei, Mei, Hongwei
Mathematics · #FOS: Mathematics #Optimization and Control (math.OC) #Probability (math.PR) #math.OC #math.PR
paper · pdf · doi:10.48550/arxiv.1909.06863
arxiv created 2020/10/21 · arxiv updated 2020/10/22
This paper is devoted to solving a time-inconsistent risk-sensitive control problem with parameter \e and its limit case (\e→0+) for countable-stated Markov decision processes (MDPs for short). Since the cost functional is time-inconsistent, it is impossible to find a global optimal strategy for both cases. Instead, for each case, we will prove the existence of time-inconstant equilibrium strategies which verify the so-called step-optimality. Moreover, we prove the convergence of \e-equilibriums and the corresponding value functions as \e→0+.